Abstract low-poly geometric background in shades of pale blue and white

Data in the Middle: The common language of research

Tools and workflows from IASSIST 2017 for research data management

The annual IASSIST conference gathered data professionals, librarians, and researchers in Lawrence, Kansas from May 23 to 26, 2017, for a programme built around the theme "Data in the Middle: The Common Language of Research." Workshops, plenaries, and lightning talks filled four days with practical demonstrations of systems that help institutions organise, describe, and preserve their research outputs. For data stewards watching from Australia, the conference offered a useful snapshot of how international peers were tackling problems familiar to repositories in Melbourne, Canberra, and Brisbane.

What made the workshop tracks particularly relevant to Australian attendees was the focus on cross-disciplinary tooling. Several presenters demonstrated platforms that work for social scientists, environmental researchers, and digital humanities scholars alike, which mirrors the multi-faculty service model common at universities such as the University of Sydney, Monash, and the Australian National University. The recap below covers the main tools and methodologies surfaced during the hands-on sessions, with notes on how they connect to initiatives already running through the Australian Research Data Commons and partner organisations.

Opening plenary and the state of data curation

The opening plenary set a tone that ran through every subsequent workshop: data curation has shifted from a back-office archival task to a frontline research skill. Speakers emphasised that modern research projects produce petabytes of mixed-format material, from genome sequences and satellite imagery to survey responses and scanned manuscripts. The volume is no longer the main obstacle; the challenge is making that material discoverable, citable, and reusable decades after collection.

This framing resonated with delegates familiar with national infrastructure projects such as the Atlas of Living Australia and the CSIRO Data Access Portal, both of which aggregate research-grade records across institutions. Several attendees noted parallels between the IASSIST discussions and the work underway at ARDC to harmonise metadata across the country's universities. The conversation made clear that curation is increasingly collaborative, with librarians, IT staff, and domain researchers sharing responsibility for the integrity of research outputs.

Software solutions demonstrated in hands-on sessions

Workshop facilitators spent considerable time on three categories of software: command-line utilities for transforming messy datasets, web-based catalogues for documenting collections, and cloud-hosted environments for analysing sensitive records. OpenRefine, R, and Python scripts featured prominently for cleaning and reshaping tabular data. Participants worked through real datasets, learning to split author names, normalise geographic fields, and convert between long and wide formats without losing provenance.

For documentation, presenters showcased platforms such as Dataverse, DSpace, and CKAN, each with different strengths depending on institutional needs. Dataverse stood out for its native support for variable-level metadata, which suits survey research projects common in social science faculties. CKAN appealed to those building public-facing open data portals, a model Australian local councils in Brisbane and Adelaide have adopted to publish transport, planning, and environmental monitoring datasets. The cloud workspaces, demonstrated using JupyterHub and Open OnDemand, illustrated how institutions can offer researchers a controlled environment without forcing them to install software locally.

Metadata standards and documentation frameworks

A persistent thread throughout the workshops was the importance of choosing metadata standards early in a project's lifecycle. Presenters walked through Dublin Core, DataCite, DDI (Data Documentation Initiative), and domain schemas such as Darwin Core for biodiversity records. The latter proved especially relevant to attendees involved in projects monitoring the Great Barrier Reef or tracking species along the Victorian coast, where Darwin Core adoption aligns with global biodiversity repositories.

Workshop leaders stressed that metadata work is iterative. Researchers rarely know all the variables they will need to describe at the start of a project, so tools that support flexible, evolving schemas save considerable effort later. Several facilitators recommended writing a data management plan at proposal stage and revisiting it annually. Australian delegates pointed out that the Australian Code for the Responsible Conduct of Research already encourages this practice, and ARDC's online DMP tool was cited as a practical resource. The sessions reinforced that documentation is not bureaucratic overhead but a foundation for reuse, replication, and trust.

Repositories, preservation, and long-term curation

Long-term preservation was a recurring concern, particularly for projects producing material of cultural significance or ongoing scientific value. Workshop facilitators described layered strategies: depositing copies in trusted repositories such as Dryad, Zenodo, or institutional archives, while keeping working copies in project-specific storage. The discussion touched on the LOCKSS framework, which replicates content across geographically distributed servers to guard against loss.

For Australian researchers, this conversation intersects directly with national investment through the National Collaborative Research Infrastructure Strategy, which funds facilities such as the Australian Antarctic Science Program and the Integrated Marine Observing System. Both rely on robust preservation pipelines to keep decades of environmental observations accessible. Workshop presenters stressed that repositories should not be treated as dump sites but as curated services with clear retention policies, version control, and withdrawal procedures when necessary.

Workflow integration and automation lessons

Several workshops focused on the unglamorous but essential work of stitching tools together into repeatable workflows. Presenters demonstrated using ORCID for researcher identification, cross-referencing funder identifiers, and harvesting metadata via OAI-PMH into discovery layers such as VIVO. Automating these steps with scripts saves hundreds of hours per year for institutional repositories, a point that resonated with small teams at regional universities in Hobart, Darwin, and Perth.

The most popular session involved building pipelines that take raw survey responses, clean them with OpenRefine, document them with a structured README, and deposit them in a repository with a persistent identifier. By the end of the workshop, participants had a working template they could adapt to their own institutions. This end-to-end approach mirrored the workflow adopted by the Melbourne Research Cloud and several faculties within the University of Queensland, where research data services have moved away from ad-hoc support towards repeatable, documented pipelines.

Applying the FAIR principles across disciplines

The FAIR principles — Findable, Accessible, Interoperable, Reusable — appeared on nearly every slide deck. Workshop leaders broke each principle into actionable steps rather than treating them as abstract ideals. Findable meant assigning DOIs and writing descriptive metadata in standard vocabularies. Accessible meant using repositories with clear licensing and authentication where needed. Interoperable meant choosing formats and schemas that other systems could read. Reusable meant rich provenance, clear licences, and links to related publications.

This framing aligned with the Maiam nayri Wingara Indigenous Data Governance principles, which Australian researchers increasingly reference when working with data collected from or about Aboriginal and Torres Strait Islander communities. Combining FAIR with culturally informed governance emerged as a strong theme, particularly in sessions led by delegates from Aotearoa New Zealand and Australian institutions. The workshops showed that the principles scale: a small qualitative study and a continent-wide environmental monitoring programme can both adhere to them with appropriate tooling.

Reflections for Australian research communities

For Australian data professionals, the IASSIST 2017 workshops offered more than a survey of available tools. They provided a chance to compare notes with international peers on persistent problems: limited staff time, uneven researcher engagement, and the challenge of preserving data across institutional reorganisations. The Australian Research Data Commons, successor to the Australian National Data Service, has invested heavily in shared infrastructure since the conference, and many of the tools demonstrated in Lawrence now underpin national services.

Delegates who attended from Australian institutions frequently reported back that the most valuable takeaway was the human network. Slack channels, mailing lists, and conference follow-up calls kept conversations alive long after the closing plenary. That community matters as much as any single platform: research data stewardship is a collective undertaking, and conferences such as IASSIST remind practitioners that they are not solving these problems alone.

Tool Primary use Strength Best fit
OpenRefine Data cleaning and transformation Handles messy tabular data with provenance Survey responses, bibliographic records
Dataverse Repository with variable-level metadata Native support for social science datasets Multi-project institutional repositories
CKAN Open data portal Strong public-facing catalogue features Government and council open data
Zenodo Long-tail preservation CERN-backed infrastructure with DOIs Small datasets, supplementary materials
JupyterHub Analytic environment Reproducible notebooks in the cloud Teaching, collaborative analysis
ORCID Researcher identification Persistent author identifiers Linking outputs to people across systems

Practical features worth adopting soon

  • Variable-level metadata support that survives schema changes
  • Persistent identifiers assigned at deposit rather than publication
  • Built-in licence selection during upload rather than after
  • Workflow templates for cleaning, documenting, and depositing in one pass
  • Community channels where curators can ask questions and share fixes

Considerations before rolling out new platforms

  • Existing infrastructure at the home institution, particularly storage capacity
  • Researcher willingness to learn a new interface
  • Long-term maintenance commitments beyond the initial project funding
  • Alignment with national initiatives such as ARDC and the FAIR principles
  • Cultural protocols required for community-owned or sensitive material

At the Conference

What attendees experienced in Lawrence

p1 reed

Plenary Sessions

Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.

platinum icpsr

Workshops

Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.

gold rockhurst

Social Events

An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.

Venue & Accommodations

Where the conference took place

silver ifdo

Kansas Union

University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.

silver ciser

The Eldridge

701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.

platinum icpsr

Springhill & TownePlace Suites

Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.

Program Highlights

Sessions and activities

Plenary Sessions Workshops Poster Presentations Committee Meetings Opening Reception Banquet Tour Kansas City Pecha Kucha Check-In Local Favorites

Getting Here

Lawrence, Kansas

Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045

Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).

Plan Your Stay

Accommodation options that were available

silver ifdo

The Eldridge

701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.

silver ciser

Springhill Suites

Marriott property. Room block reserved under "KU IASSIST Conference."

gold rockhurst

TownePlace Suites

Marriott property. Room block reserved under "KU IASSIST Conference."