Disclosure Control of neurominority groups – Protection or Erasure?

Posted on

Data protection is important – so what’s the problem?

Written by Cara Kendal

Statistical Disclosure Control is an important process where statistical data is protected before being released to make sure that no identifying information is made publicly available. This is a large task that is often completed by several people. One of the primary methods of SDC is output checking. This is when outputs from data (tables, graphs, statistics etc.) are observed both by the initial researcher producing the outputs and a pair of output checkers associated with the data holders to ensure that the data can’t be extracted from the outputs in a way that would compromise anonymity.

Although there are rules that an output checker will look for when attempting to secure data, perhaps the most common of these is count thresholds; a measure that places a minimum limit on the number of observations that can be present in any cell of a table or count on a graph. This prevents someone who might have some knowledge of a participant in a study from working out the rest of their data. Typically, data holders will have a specific value below which data requires suppression.

Usually, this is enough to protect data with only a slight reduction in detail level. This can become problematic, however, when collected data refers to a minority demographic.

For example, in a hypothetical study, the researcher is surveying employees in order to observe career satisfaction. Of 500 employees surveyed, only seven identified as disabled, two of which stated that they are diagnosed with an Autism Spectrum Condition. By the rules of SDC, these results would need to be suppressed to protect the identities of those involved. However, when observing the data, it is clear that those participants who identified as disabled have a large reduction in work satisfaction compared to their peers The researcher wants to use these results to justify further investigation into this issue in the form of a large survey, specifically targeting this group but would struggle to support this properly in their paper while still concealing the suppressed data points.

This hypothetical demonstrates a difficult scenario where there is no clear answer. Suppressing the data would prevent any risk of disclosure of confidential health data but would potentially prevent to furthering of research into an under-represented group.

Data as a weapon

An argument for stricter protection of data is the potential harm that data can cause when weaponised against minority groups (Jan Trust, 2021). It is known that Data can have a potential risk when released, especially when specific minorities could be targeted by that data. This topic hit the headlines in July 2020 when postcode-level testing data was withheld from publication. There was concern amongst local councils that the higher prevalence of the COVID-19 virus in BAME communities would incentivise racially motivated discrimination against these groups. This line of thought would add caution to the collection and release of minority data in an attempt to prevent the fuelling of pre-existing stigmas.

Forbes defines four indications of Data weaponization (Murphy, 2020), all of which may have real consequences when used against minority groups.

1.) Creating Selfish Proof – The cherry-picking of data to serve as confirmation bias.

2.) The Dangers of Short-Termism – prioritising immediate results over long-term consequences.

3.) The Art of Manipulation – intentionally using data to profit in a way that harms the consumer.

4.) Reducing people to their lowest common denominator – defining a person by a behaviour or action out of context.

It‘s not difficult to see how all of these could be dangerous, especially when aimed at vulnerable individuals. So, when considering neurodiverse populations, many of whom face immense stigma in their day-to-day lives, it is important to turn a critical eye to what drawing attention to this group will achieve. Is this an area where neurodiverse people may have widely different experiences from the neurotypical population? Is this an area where neurodiverse voices have been historically undervalued? Or conversely, will singling this group out in particular unnecessarily place them at risk of data weaponization?

Data as a tool of representation

Accurate and detailed data about minority groups is vital to properly understanding the needs and experiences of those groups as well as making sure that they have space to fit into the tapestry of  larger population observations.

A recent example of this is the questions surrounding sexuality and gender in the 2021 UK census(Moss, 2023).. For the first time in history the UK Census simply included one question on how the participant would define their sexuality, one on if their gender identity was the same as their sex assigned at birth (with the option to write in any non-cis-gendered identity) and one question asking if a married individual was married (or in a civil partnership) with a member of the same or opposite sex. While brief, and perhaps a little surface-level, this line of questioning provided an accurate look at the population level for the first time. Because of  the collection of this data, areas with higher levels of LGBTQIA+ individuals can be identified, demonstrating the need for leaders of those local authorities to hold their LGBTQIA+ constituents in mind when acting in their role. While useful in gaining a basic overview of the number and location of queer populations in the UK, this data set does not gather much detail on the nuances of identity within these communities.

In comparison, the Gender Census , which attracted 39,765 participants in 2022, collects more specific data and a different approach to disclosure control. While the data holder does produce a full report of the data and trends observed when compared to previous censuses each year, the full data set is made available on the website. While each data entry is associated with a random participant number rather than any name or IP address, the survey does not apply cell suppression to any given information other than the country of origin (with a minimum count of ten). This allows data pieces with only one or two entries to remain in the public data set, acknowledging lesser-used identity words or pronoun sets and providing  a more complete landscape of the language used by this community to describe themselves. However, this does mean that the issues that would be managed by output checking remains; a bad actor could find a participants full data set by searching for a known, rare response the participant had given. In this case the researcher addresses this by being explicit about how and why the data collected is stored and distributed and giving participants the option to have their data removed at any time if they are no longer comfortable with being involved in the research.

On the website for this survey, they have a tab dedicated to data protection where they speak in detail about their data procedures and their limitations. By acknowledging this openly, they are able to collect the much needed and high-quality data covering a large participant group.

The Case of Neurodiversity – An Autistic Example

When considering these issues in connection with neurodiverse populations, it’s important to acknowledge the history of exclusion this group has experienced within research. The autistic community have, for example, long been isolated from the researchers investigating their disorder (Milton, 2019). Autistic researchers and advocacy groups have been pushing for direct involvement in research as many autistic people (academics included) feel that there is a real disparity between current research and research that would actively be useful to the community (Loughran, 2020). Because of this, many autistic advocacy groups have adopted the slogan ‘Nothing About Us Without Us’.

Throughout both historic and current research, there has been a huge focus on developing a ‘cure’ for Autism Spectrum Conditions. The fact that the majority of the autistic community do not want a cure as they feel that it is impossible to separate their autism from who they are, a state of being they don’t believe is ‘wrong’ or requires any correcting has been ignored.

This primarily leads to two responses within the community: Autistic individuals who are resistant to and distrustful of autism research out of fear that their participation will lead to an attempt to un-consensually ‘fix’ them, and autistic people (often academics and researchers themselves) who actively fight to advocate for themselves and their community. They wear the label autistic with pride and seek to educate those around them as well as those in academia on what it is like to live as an autistic person.

Research into autism is needed desperately in many areas, from healthcare to education to employment. As of 2021, only 21.7% of autistic people were in employment (Cusack, 2021), the lowest of any disabled group. Further research into why this is the case could trigger social change that would benefit hundreds of thousands of autistic adults in the UK and have a positive rippling effect into UK society as a whole.

However, in order to help groups like these it is important that data collected on them is as detailed as possible and their data isn’t disregarded, even when they only appear in small numbers.

Solutions?

There is no one answer to this problem. The issue is complex and commands many strong, apposing arguments. This most appropriate response to this concern will likely differ on a case-by-case basis. However, here are two solutions that may be useful.

One could combine labels referring to specific conditions into larger group labels. Such as ‘Neurodivergent’, ‘Neurological conditions’, etc. This would allow more data sets to be grouped in a way that would aid in the prevention of data disclosure without removing the distinction entirely. This may be useful in some cases however has the risk of severally reducing the specificity of the research. By combining these labels into one larger category, it incorrectly assumes  the individual groups within it have synonymous experiences and needs. It also has the downside of erasing the individual identities, many neurodivergent individuals take pride in their specific labels and seek to reduce stigmatisation around that language. This mirrors the conversation currently happening around the BAME acronym (UWE Bristol, 2022).

Another possible solution would be to collect specific data with direct informed consent from those involved. By having consent discussions with participants, especially those known to the researcher to be of a minority group, with honest explanations of the risk and granting them the opportunity to rescind participation in the research should they be uncomfortable with that risk, the researcher is able to collect and properly analyse the data on a minute scale. This would not remove the risk of disclosure, but would allow participants to be involved in the discussions around how to handle their data.This could lead to participants workshoping ideas with the research team to allow their data to be used in a way that is both specific and protective. A potential downside of this method is that it may deter some potential participants from being willing to be involved in the research, thus having a detriment to population size.

There may well be many other ways to face this issue, if you have any suggestions or queries, I would love to hear from you. You can reach me at: cara.kendal@uwe.ac.uk


References

Aspinall, P. J. (2014). Identifying key vulnerable groups in data collections: vulnerable migrants, gypsies and travelers, homeless people, and sex workers. Center for Health Services Studies, University of Kent, 18.

Cusack, J. (2021, February 19). Autistic people still face highest rates of unemployment of all disabled groups. Retrieved March 1, 2023, from https://www.autistica.org.uk/news/autistic-people-highest-unemployment-rates

How can data be weaponised to target marginalised groups? (2021, April 06). Retrieved February 28, 2023, from https://jantrust.org/blog/how-can-data-be-weaponised-to-target-marginalised-groups/

Kapadia, D. (2021, February 18). Represented yet excluded: How ethnic minority people are counted in national surveys. Retrieved March 1, 2023, from https://blog.ukdataservice.ac.uk/represented-ethnic-minority-people/

Loughran, E. (2020, February 10). Why autism research needs more input from autistic people. Retrieved March 1, 2023, from https://www.spectrumnews.org/opinion/viewpoint/why-autism-research-needs-more-input-from-autistic-people/

Milton, D. (2014). What is meant by participation and inclusion, and why it can be difficult to achieve.

Milton, D. E. (2019, August 15). Beyond tokenism: Autistic people in autism research. Retrieved March 1, 2023, from https://www.bps.org.uk/psychologist/beyond-tokenism-autistic-people-autism-research

Moss, L. (2023, January 06). Census data reveals LGBT+ populations for first time. Retrieved March 1, 2023, from https://www.bbc.co.uk/news/uk-64184736

Murphy, B. (2020, October 27). Council post: Weaponizing data versus using it as a tool to humanize consumers. Retrieved March 1, 2023, from https://www.forbes.com/sites/forbesagencycouncil/2020/10/28/weaponizing-data-versus-using-it-as-a-tool-to-humanize-consumers/?sh=635f0c022f04

NB/GQ Survey 2022 – the worldwide results. 2022, August 22. https://www.gendercensus.com/results/2022-worldwide/

Singh, A. (2020, July 14). Exclusive: Covid Test Data held back from publication over community cohesion concerns. Retrieved February 28, 2023, from https://www.huffingtonpost.co.uk/entry/coronavirus-testing-data-councils-government_uk_5f0de1bdc5b648c301f02d00?9s

United Nations. (2022, February 16). Better data collection bolsters human rights of marginalised people. Retrieved March 1, 2023, from https://www.ohchr.org/en/stories/2022/02/better-data-collection-bolsters-human-rights-marginalised-people

UWE Bristol. (2022, January 21). University drops use of BAME acronym. Retrieved March 1, 2023, from https://info.uwe.ac.uk/news/uwenews/news.aspx?id=4206

Leave a Reply

Your email address will not be published. Required fields are marked *

Back to top