Lydia Lucchesi (ANU) – Smallset Timelines: A Visual Representation of Data Preprocessing Decisions

Date & Time:

July 18, 2022 3:00 pm – 4:00 pm

Location:

Crerar 346, 5730 S. Ellis Ave., Chicago, IL,

07/18/2022 03:00 PM 07/18/2022 04:00 PM America/Chicago Lydia Lucchesi (ANU) – Smallset Timelines: A Visual Representation of Data Preprocessing Decisions UChicago HCI Club Seminar Crerar 346, 5730 S. Ellis Ave., Chicago, IL,

Data preprocessing is a crucial stage in the data analysis pipeline, with both technical and social aspects to consider. Yet, the attention it receives is often lacking in research practice and dissemination. We present the Smallset Timeline, a visualisation to help reflect on and communicate data preprocessing decisions. A “Smallset” is a small selection of rows from the original dataset containing instances of dataset alterations. The Timeline is comprised of Smallset snapshots representing different points in the preprocessing stage and captions to describe the alterations visualised at each point. Edits, additions, and deletions to the dataset are highlighted with colour. We develop the R software package, smallsets, that can create Smallset Timelines from R and Python data preprocessing scripts. Constructing the figure asks practitioners to reflect on and revise decisions as necessary, while sharing it aims to make the process accessible to a diverse range of audiences. We present two case studies to illustrate use of the Smallset Timeline for visualising preprocessing decisions. Case studies include software defect data and income survey benchmark data, in which preprocessing affects levels of data loss and group fairness in prediction tasks, respectively. We envision Smallset Timelines as a go-to data provenance tool, enabling better documentation and communication of preprocessing tasks at large.

Speakers

Lydia Lucchesi

PhD Student, Australia National University

Lydia is a PhD Candidate in Computer Science at the Australian National University. She completed a BA in statistics at the University of Missouri, USA, followed by a post-bachelor fellowship at the Institute for Health Metrics and Evaluation. Her current research focuses on the visualisation of data quality. She is a co-developer of the Vizumap R package, a toolkit for visualising uncertainty in spatial data.

Resources

Community

What’s Real and What’s Not? Watermarking to Identify AI-Generated Text

Enhancing Multitasking Efficiency: The Role of Muscle Stimulation in Reducing Mental Workload

From wildfires to bird calls: Sage redefines environmental monitoring

“Machine Learning Foundations Accelerate Innovation and Promote Trustworthiness” by Rebecca Willett

Nightshade: Data Poisoning to Fight Generative AI with Ben Zhao

Ian Foster – Better Information Faster: Programming the Continuum

Speakers

Lydia Lucchesi

Unveiling Attention Receipts: Tangible Reflections on Digital Consumption

Five UChicago CS students named to Siebel Scholars Class of 2024

UChicago Computer Scientists Design Small Backpack That Mimics Big Sensations

UChicago Team Wins The NIH Long COVID Computational Challenge

UChicago Assistant Professor Raul Castro Fernandez Receives 2023 ACM SIGMOD Test-of-Time Award

Computer Science Class Shows Students How To Successfully Create Circuit Boards Without Engineering Experience

UChicago CS Researchers Shine at CHI 2023 with 12 Papers and Multiple Awards

New Prototypes AeroRigUI and ThrowIO Take Spatial Interaction to New Heights – Literally

PhD Student Kevin Bryson Receives NSF Graduate Research Fellowship to Create Equitable Algorithmic Data Tools

Computer Science Displays Catch Attention at MSI’s Annual Robot Block Party

UChicago / School of the Art Institute Class Uses Art to Highlight Data Privacy Dangers

UChicago, Stanford Researchers Explore How Robots and Computers Can Help Strangers Have Meaningful In-Person Conversations