Blog post #1. Defining data and its distribution.

Data is not as unbiased as one would think it is. How we collect data is essential on how it will be processed, interpreted and used in a project that can affect millions. Raw data in its essence can be considered a nonpartisan entity but the method it was collected can be biased. Ranging from collecting data for commercial use and all the way to using it for policing members of the community, The “humanities” in collecting data is often overlooked and not criticized enough  

For example, in Introduction by Jacqueline Wernimont and Elizabeth Losh, we can view their strong opinion that Digital Humanities should not have a linear methodology in collecting and presenting data, the world is broad, and data must have a tie to multiple sections. Inserting a feminist lens can aid in introducing intersectionality into the field. Jacqueline and Elizabeth state that” our argument is that feminisms have been and must continue to be central to the identity and the methodologies of the digital humanities as a field.” intersectionality in this field can be a huge benefit as we are collecting data from real life scenarios.  

In order to properly introduce intersectionality, first we should ask the question of “Should there be more training with data collection?” We can see that there are various forms of methodology in metadata collection, some come with biases and others try to correct them. We see again, in Introduction, the authors state, “In addition to adding new vocabulary to existing taxonomical systems, they assert that queerness also points toward a shift in the very methodologies of metadata collection. To queer metadata, queer thinking must be brought to bear on the conceptual models and tools of object description to challenge the norms that dictate how meaning is derived from data. They observe that the methods with which data are traditionally mapped rely on a model of the one-to-one relationship between concepts of the world that can account for nonbinary relationships.” While there are many data humanists who are allies to these taxonomical systems, we can assume that there are just as many that do not understand these terms and often find ways and, unwillingly, cause some misinterpretation in data collection. Often not due to malice but due to ignorance. A simple example can be found in the American census system in which many demographics are not properly counted, and this can lead to changes that can affect millions.  

There can be tension in trying to bring this information forward to the overall field.  I agree with the interpretation of Beth Coleman’s response to how data is used against members in the community when the 2020 protest started.  People’s livelihoods were at stake when authority figures started to collect large sets of data and then started narratives such as grouping a demographic into one pot and demonizing them in the media. I personally feel that the Digital humanities needs to have a standard on collecting data which I understand can be difficult.  

References and readings:

Jacqueline Wernimont and Elizabeth Losh. 2018. “Introduction” In Bodies of Information: Intersectional Feminism and Digital Humanities, edited by Jacqueline Wernimont and Elizabeth Losh. University of Minnesota Press.

Gold, Matthew K. 2012. The Digital Humanities Moment In Debates in the Digital Humanities, edited by Matthew K. Gold. University of Minnesota Press.

1 thought on “Blog post #1. Defining data and its distribution.

  1. Caitlin Cacciatore (she/hers)

    I agree that data bias is a huge issue in an increasingly automated world.

    Some examples that I always return to are less far-reaching than your example of the census, but just as insidious: such as the fact that the majority of fitness trackers on the market do not work as well on darker skin – simply because of the data points collected during testing, which sampled lighter, whiter skin much more than it sampled darker, more melanated skin. A fly in the ointment at a stage that early in the development gums the entire work.

    This is part of a larger system of oppression and microaggressions and infrastructure that discriminates, consciously or not, against the BIPOC community. Neighborhoods that are predominately African American are more likely to be considered food deserts, and the highest rates of obesity are among African American women, according to the US Department of Health and Human Services Office of Minority Health. And yet, fitness trackers don’t work as well on darker skin – a vicious cycle that traps BIPOC individuals in a relentless cycle of bodily harm and lower life expectancies.

    Another example I turn to often is the fact that bank loans were historically not given to BIPOC individuals nearly as often as they were given to people of other demographics. Fed historical data, the automated systems used to weed people out who were requested loans recognized this through a combination of zip codes with a higher percentage of African Americans, as well as other data that might seem innocuous and unrelated to race. This perpetuates a system of discrimination.

    Ignorance, in my opinion, is no excuse. LIves and livelihoods are at stake. There must indeed be a system by which data becomes more inclusive and equitable, and the time to act upon such a standard is now.

Comments are closed.