Enabling Responsible Data Science through Multi-Dimensional Data Management

Lin, Yin

Enabling Responsible Data Science through Multi-Dimensional Data Management

dc.contributor.author	Lin, Yin
dc.date.accessioned	2025-05-12T17:34:54Z
dc.date.available	2025-05-12T17:34:54Z
dc.date.issued	2025
dc.date.submitted	2025
dc.identifier.uri	https://hdl.handle.net/2027.42/197083
dc.description.abstract	In today’s world, data is collected and utilized at an unprecedented scale, profoundly influencing society. Big data enables analyses that guide high stakes decisions and supports data-driven systems. While the benefits of big data are significant, the challenges extend beyond efficient processing and storage; as data scientists, we are responsible for ensuring that data applications ethically benefit society. Data science technologies can cause harm if they reinforce inequities, particularly when sensitive data—such as data linked to protected characteristics like race and gender is mishandled. This dissertation contributes to responsible data science by proposing comprehensive data management techniques to address challenges throughout the big data lifecycle. These approaches aim to enhance fairness, transparency, and accountability in data systems while considering the complexities associated with multiple protected characteristics. Firstly, data acquisition often results in the underrepresentation of certain populations, which risks perpetuating unfair treatment and oversight of these groups. Obtaining representative samples becomes particularly challenging when dealing with intersectional subgroups. We propose coverage analysis techniques to efficiently identify representation bias in multi-table databases, guiding data users toward obtaining more representative samples. Secondly, biases embedded in historical decisions can propagate into downstream machine learning tasks, resulting in unfair predictions. We emphasize the importance of addressing the root causes of unfairness in the training data. We propose model agnostic data pre-processing techniques to effectively detect and mitigate biased data collection, thereby enhancing ma- chine learning fairness across subgroups. Thirdly, data analytics based on cherry-pick generalizations can lead to misleading insights, diminishing the experiences of certain subgroups in decision-making. We refine these generalizations across multiple attributes to develop a framework evaluating their appropriateness, identify subgroup discrepancies, and promote more accurate and inclusive representations of data. Lastly, when a data analysis pipeline produces unexpected outputs, it is the responsibility of data scientists to interpret the potential sources of error. We propose a row-level data lineage approach to enhance pipeline transparency, enabling them to trace the origins of issues.
dc.language.iso	en_US
dc.subject	Responsible Data Science
dc.subject	Data Management
dc.subject	Database
dc.subject	AI fairness
dc.title	Enabling Responsible Data Science through Multi-Dimensional Data Management
dc.type	Thesis
dc.description.thesisdegreename	PhD
dc.description.thesisdegreediscipline	Computer Science & Engineering
dc.description.thesisdegreegrantor	University of Michigan, Horace H. Rackham School of Graduate Studies
dc.contributor.committeemember	Jagadish, H V
dc.contributor.committeemember	Garcia, Patricia
dc.contributor.committeemember	Fish, Benjamin
dc.contributor.committeemember	Mozafari, Barzan
dc.subject.hlbsecondlevel	Computer Science
dc.subject.hlbtoplevel	Engineering
dc.contributor.affiliationumcampus	Ann Arbor
dc.description.bitstreamurl	http://deepblue.lib.umich.edu/bitstream/2027.42/197083/1/irenelin_1.pdf
dc.identifier.doi	https://dx.doi.org/10.7302/25509
dc.identifier.orcid	0000-0002-6609-5706
dc.identifier.name-orcid	Lin, Yin; 0000-0002-6609-5706	en_US
dc.working.doi	10.7302/25509	en
dc.owningcollname	Dissertations and Theses (Ph.D. and Master's)

Files in this item

Name:: irenelin_1.pdf
Size:: 5.645MB
Format:: PDF

View/Open

Dissertations and Theses (Ph.D. and Master's)

Show simple item record

Remediation of Harmful Language

The University of Michigan Library aims to describe its collections in a way that respects the people and communities who create, use, and are represented in them. We encourage you to Contact Us anonymously if you encounter harmful or problematic language in catalog records or finding aids. More information about our policies and practices is available at Remediation of Harmful Language.

Accessibility

If you are unable to use this file in its current format, please select the Contact Us link and we can modify it to make it more accessible to you.