Practical Examples from Data in the Middle at IASSIST 2017
The IASSIST 2017 conference gathered data professionals, archivists, librarians and researchers at the University of Kansas in Lawrence between 23 and 26 May 2017. The theme "Data in the Middle: The Common Language of Research" reflected a growing awareness that the people who curate, clean, manage and translate data sit at the very centre of contemporary scholarship. Across plenary sessions, workshops and poster presentations, contributors shared tangible workflows and lessons that participants could take back to their own institutions.
For readers based in Australia, the Kansas discussions carry particular weight. The country faces long distances between researchers, a research data landscape coordinated through the Australian Research Data Commons, and unique tropical and Indigenous datasets that demand careful stewardship. Practical examples from Lawrence therefore offer useful parallels for anyone working with the Australian Bureau of Statistics, university research offices or partnerships such as Data61.
Big Data Workflows in Practice
Several Lawrence presentations tackled what actually happens once a research group has terabytes of raw material to wrangle. One recurring example came from climate and earth science teams, who described end-to-end pipelines built around containerised scripts and version-controlled storage. Instead of emailing enormous CSV attachments around, teams used shared file systems and reproducible build environments so that every collaborator saw identical inputs and outputs.
Another workflow highlighted at the conference focused on social science survey data. Speakers from European statistical agencies demonstrated how cleaning, anonymising and documenting survey responses could be split into discrete steps, each logged in a run sheet. This made it possible to audit transformations months later, which is increasingly demanded by funding bodies. Australian readers running longitudinal studies through the Australian Bureau of Statistics or HILDA will recognise the same pressure to show provenance.
The most striking practical lesson was the importance of naming conventions. Researchers who once used cryptic folder labels like "final_v3_REAL" reported moving to date-stamped, project-coded structures. As one presenter put it, the data managers in the middle are the ones who save a project when a postdoc leaves unexpectedly.
Deep Learning and Data Preparation
The deep learning sessions at IASSIST 2017 focused less on flashy neural network results and more on the unglamorous work that makes those results possible. Several presenters described image classification projects where most of the effort went into labelling thousands of training examples accurately. Teams using citizen-science platforms reported needing robust feedback loops, because volunteer labels tend to drift if quality control is missing.
A second theme involved the use of pre-trained models for text and document analysis. Librarians from large university systems explained how they applied off-the-shelf language models to draft descriptive metadata for digital collections, then edited the output by hand. This hybrid approach, blending machine suggestion with human review, is gaining traction among Australian cultural heritage institutions that hold vast un-catalogued archives of Pacific and Asian material.
Conference participants also discussed ethical questions around training data. Presenters stressed that researchers must document the provenance of any corpus used to train a model, particularly when it contains personal information. For Australian researchers, this connects directly to discussions about Indigenous data sovereignty and the principles developed by groups like the Maiam nayri Wingara Aboriginal and Torres Strait Islander Data Sovereignty Collective.
Digital Agriculture and Global Food Security
Digital agriculture emerged as a headline topic at Lawrence, with sessions running across two of the four days. Speakers from international agricultural research centres showed how sensor data from soil moisture probes, satellite imagery and yield monitors can be combined to advise smallholder farmers in real time. The key practical lesson was interoperability: equipment from different vendors rarely speaks the same data language, so middleware layers are essential.
A workshop on food security data emphasised the value of standardised crop and weather variables. Researchers walking through actual case studies from sub-Saharan Africa and South-East Asia demonstrated how missing or misaligned metadata can render cross-country comparisons meaningless. Conference attendees from the Australian Centre for International Agricultural Research found the parallels with their own regional work in the Indo-Pacific immediately useful.
The conference showcased practical visualisation tools for agricultural dashboards. Presenters stressed that dashboards only earn the trust of farmers and policy makers when the underlying data lineage is clear. Anyone interested in further applied data work can explore additional case studies at MATIS Projects, where similar middle-layer data coordination challenges are documented.
Reproducible Research as a Working Standard
Reproducibility was a thread running through nearly every session, but a dedicated strand pulled it into practical focus. Presenters described moving from reproducibility as an aspiration to reproducibility as a default workflow. One university team shared their internal checklist, which included requirements around project registration, pre-analysis plans and shared code repositories.
The Australian context came up several times. The Australian Research Data Commons has invested heavily in platforms that support reproducible workflows, and conference attendees from institutions like the University of Sydney, the University of Melbourne and the Australian National University shared how they had adopted similar templates locally. In particular, attendees noted how having a yarn with colleagues over a coffee in the staff kitchen often surfaced the most useful reproducibility tips, a reflection of how relaxed Australian workplace culture can accelerate informal knowledge sharing.
A practical example came from a Queensland-based group presenting at Lawrence. They described how they had embedded a data management plan into the ethics approval process itself. By making the plan a precondition for ethics clearance, they avoided the all-too-common situation where data ends up scattered across personal laptops at the end of a project.
Teaching Data Literacy Across Disciplines
Education and training sessions explored how data skills are taught to students who do not see themselves as statisticians or programmers. Conference presenters from the United States, Canada and Europe shared curriculum models that begin with students' own research questions and introduce tools only as needed. This contrasts with older approaches that started with software tutorials and hoped students would later apply them.
Australian participants drew comparisons with their own teaching contexts. Regional universities in places like Wollongong, Hobart and Darwin often work with smaller cohorts and need flexible teaching materials. Several conference slides featured free, openly licensed resources that can be adapted without licensing fees, a critical factor for institutions with tight training budgets. The conference's tone was collegial rather than prescriptive, more like a chat over a flat white than a formal lecture.
One particularly lively session involved early-career researchers presenting their first attempts at building teaching modules on data ethics. They found that students responded strongly to case studies drawn from local controversies, such as the Robodebt scheme, which highlighted the real consequences of poor data practices in government. This grounded the abstract notion of data ethics in lived Australian experience.
Building Sustainable Data Communities
The closing sessions turned to the long-term sustainability of data services. Conference speakers warned that one-off projects, however well-funded, rarely leave lasting infrastructure behind. Instead, they advocated for embedded data stewards within research units, supported by institutional funding rather than short-term grants.
Practical examples came from collaborative consortia. One North American consortium described pooling a fraction of every grant it received into a shared fund for data infrastructure. The model has since spread to other regions, and Australian attendees from the Australian Research Data Commons expressed interest in adapting it to fund national data assets. The conversation acknowledged the particular challenge of serving researchers spread across enormous distances, from Cairns to Perth.
A final panel reflected on the conference theme itself. The middle layer of data work, often invisible in research outputs, deserves recognition, support and stable career paths. Speakers urged participants to advocate within their own institutions for data professionals to be named as co-investigators and acknowledged as authors on papers that depend heavily on their work.
Comparing Approaches Across the Conference
Different presentations at IASSIST 2017 took different routes to similar practical goals. The comparison below summarises how a few key sessions approached workflow, reproducibility and audience.
| Session Theme | Workflow Approach | Reproducibility Emphasis | Primary Audience |
|---|---|---|---|
| Big data pipelines | Containerised scripts and shared storage | Versioned builds and run sheets | Climate and survey researchers |
| Deep learning prep | Hybrid human and machine labelling | Documented training corpora | Librarians and digital humanities teams |
| Digital agriculture | Sensor integration and middleware | Standardised crop and weather variables | Agricultural researchers and policy makers |
| Reproducible research | Embedded data management plans | Pre-registration and ethics-linked plans | University research offices |
| Data literacy teaching | Question-led curriculum design | Open educational resources | Lecturers and librarians |
Practical Takeaways for Australian Data Practitioners
Several lessons from Lawrence translate directly into the Australian setting, particularly for those working outside the major east-coast research hubs.
- Adopt project-coded, date-stamped folder structures instead of informal version labels
- Embed data management plans inside ethics approval processes to prevent data loss
- Use hybrid labelling workflows when applying machine learning to cultural collections
- Standardise variable names before attempting cross-country agricultural comparisons
- Allocate a small percentage of every grant to a shared data infrastructure fund
- Encourage informal knowledge sharing sessions, since the Aussie tradition of a quick yarn over coffee remains a powerful way to spread good practice
Local Initiatives Worth Watching
Australia already hosts several initiatives that align with the conference's practical themes and offer ready-made examples for new projects.
- The Australian Research Data Commons national platforms, which support reproducible workflows across universities
- The Australian Bureau of Statistics Data by Region tool, which exemplifies transparent public sector data delivery
- AURIN, the Australian Urban Research Infrastructure Network, supporting urban and infrastructure researchers
- Data61, CSIRO's data science arm, working on applied machine learning for government and industry
- The Maiam nayri Wingara Indigenous Data Sovereignty Collective, leading discussions on culturally appropriate data stewardship
At the Conference
What attendees experienced in Lawrence
Plenary Sessions
Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.
Workshops
Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.
Social Events
An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.
Venue & Accommodations
Where the conference took place
Kansas Union
University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.
The Oread
1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.
The Eldridge
701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.
Springhill & TownePlace Suites
Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.
Program Highlights
Sessions and activities
Getting Here
Lawrence, Kansas
Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045
Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).
Plan Your Stay
Accommodation options that were available
The Eldridge
701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.
The Oread
1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.
Springhill Suites
Marriott property. Room block reserved under "KU IASSIST Conference."
TownePlace Suites
Marriott property. Room block reserved under "KU IASSIST Conference."