Seeking Anonymized Field Data Collection Datasets for an Open Benchmark

Hello Admin, Not sure if this is the right place to post thi, if its not im happy to move or delete it.

Hi everyone,

I’m working on an initiative to create an open benchmark dataset for field data quality assurance.

Today, there are many excellent digital data collection platforms—such as KoboToolbox, SurveyCTO, ODK, CommCare, Survey Solutions, CSPro, and others—but there are very few publicly available datasets that developers and researchers can use to evaluate field data quality tools.

I’m looking for organizations that may be willing to share completed, fully anonymized datasets from field data collection projects, where they have the necessary permissions to do so.

I’m especially interested in datasets that include:

  • GPS coordinates (or generalized locations)

  • Interview photos

  • Audio recordings

  • Interview start and end times

  • Submission timestamps

  • Enumerator IDs (anonymized)

  • Supervisor review outcomes or quality flags (if available)

These datasets will help create a community benchmark for testing quality assurance methods such as:

  • GPS verification

  • Duplicate image detection

  • Audio quality assessment

  • Interview duration analysis

  • Duplicate submission detection

  • Fieldwork anomaly detection

The objective is to create a resource that benefits researchers, NGOs, software developers, and the wider field data collection community by making it easier to evaluate and improve quality assurance tools.

If your organization has a completed project that could be shared in an anonymized form—or if you know of existing public datasets—I would greatly appreciate hearing from you.

I’m also happy to discuss data-sharing agreements, attribution, licensing, or any requirements needed to ensure the data is used responsibly.

Thank you!

@bibiladeoyeleke

This is definitely a very interesting topic to explore. With the use of Generative AI, it’s absolutely possible to create high‑quality synthetic datasets, and several organizations are already doing work in this area.

Here are examples that might be useful for the discussion:

1 Like

Thank you so much for this. I checked and i dont think they have images and audio, do you maybe know of any other resouce that have those. Im really grateful you chimed in. Cheers.

Hi @bibiladeoyeleke! While real image and audio files could be difficult to obtain synthetically, Kobo’s research team recently worked on a project where we generated synthetic audio transcripts to assess the effectiveness of AI-assisted qualitative analysis (see report here – p8 on Synthetic Transcript Corpus). This could be an approach to consider for your work in case you are not able to obtain actual audio files from other organizations.