Abstract low-poly geometric background in shades of pale blue and white

Data in the Middle: The common language of research

What IASSIST 2017 Taught Us About Data Documentation

IASSIST 2017 gathered researchers, librarians, archivists and data specialists around a practical question: how can data move between people and disciplines without losing its meaning? Held at the University of Kansas in Lawrence from 23–26 May 2017, the conference used the phrase “Data in the Middle” to describe data as a shared language rather than a finished product.

The archived programme reflects a moment when big data, deep learning, digital agriculture and global food security were changing research practice. These fields needed large, complex and frequently reused datasets. They also exposed a basic weakness: data can be technically available while remaining difficult to interpret, assess or reuse.

Documentation was therefore treated as part of the research infrastructure. A dataset required more than a file name and a download link. Users needed to know who created it, how observations were collected, what terms and codes meant, which transformations had been applied, and where uncertainty remained.

That lesson remains highly relevant in Australia. A dataset held by a university in Canberra, a government agency in Melbourne or a research station near Dubbo may be used years later by someone in Perth or Darwin. Clear metadata, provenance and responsible access conditions help those users understand the material without relying on an informal handover or a hurried “no worries, it’s all in the folder” explanation.

Documentation Approach What It Captures Why It Matters for Reuse
Basic file description Title, creator, date and format Helps users identify the resource
Structured metadata Variables, methods, units and subject terms Makes data searchable and interpretable
Provenance record Processing steps, versions and sources Supports verification and reproducibility
Contextual documentation Purpose, population, limitations and decisions Prevents misleading analysis
Access and governance notes Consent, licences, restrictions and cultural obligations Enables ethical and lawful reuse

Data Needs Meaning Before It Needs Scale

The conference themes showed why documentation becomes more important as datasets grow. A small spreadsheet may be understood by the person who created it, but a national sensor network, satellite archive or machine-learning training set cannot depend on personal memory. Scale separates data from the circumstances that produced it.

Big data projects often combine records from different systems. A variable called “region” could mean a postcode, a statistical area, a farm district or an administrative boundary. A date may refer to collection, publication, processing or an event itself. Without a data dictionary and clear definitions, technically correct analysis can produce an incorrect interpretation.

The same issue appears in Australian research. A digital agriculture project around Wagga Wagga might combine soil readings, satellite imagery, rainfall records and machinery logs. Each source could use different units, time zones, spatial resolutions and missing-value codes. Documentation creates the bridge between those sources, allowing a later analyst to understand what was joined and what assumptions shaped the result.

This is why data documentation should be considered part of the dataset rather than an optional attachment. A useful record explains the research question, the population or environment observed, the collection instruments, the sampling design and the boundaries of valid use. It gives data a readable social and scientific context.

Metadata Is A Working Language

Metadata is sometimes treated as an administrative requirement, especially when a repository asks for fields to be completed before publication. IASSIST 2017 pointed towards a richer view: metadata is a working language between disciplines. It allows a social scientist, statistician, librarian, software developer or policy analyst to recognise what a resource contains and whether it is suitable for their purpose.

Good metadata operates at several levels. Discovery metadata helps someone find the resource. Structural metadata explains how files and variables relate. Descriptive metadata records methods and content. Administrative metadata covers ownership, licences, preservation and access. Each layer answers a different question, and a strong data record connects them.

For Australian institutions, this can mean aligning local practice with repository and catalogue systems while preserving the language used by the research community. A university might deposit survey data in its institutional repository, register a collection through the Australian Research Data Commons, and link related publications through a persistent identifier. Those systems are useful only when the underlying descriptions are consistent.

A practical metadata record should include:

  • A precise title, creator list, affiliation and date range
  • A plain-English abstract describing the data and its research purpose
  • Variable definitions, units, coding schemes and controlled vocabularies
  • Geographic coverage, time period, sampling approach and population
  • File formats, software requirements, version information and checksums
  • Licence details, contact points, access conditions and known limitations

Metadata also needs an audience beyond specialists. Researchers may understand “NDVI” or “hierarchical linear model”, while community partners, journalists and policy staff may not. A short plain-language description alongside technical detail makes a collection more useful without weakening its precision.

Provenance Makes Research Traceable

Documentation becomes especially valuable when data changes over time. Cleaning, recoding, merging and anonymising are ordinary research activities, yet they can make a published file differ substantially from the original observations. Provenance records show what happened between collection and release.

A reliable provenance trail identifies source files, software, scripts, transformations, responsible people and dates. It may link a raw dataset to an intermediate file and then to the version used in a publication. Version numbers and persistent identifiers help users distinguish an amended release from an earlier one. This is vital when a result needs to be checked months or years later.

The deep-learning discussions associated with the conference themes make this point particularly clear. A model is shaped by its training data, preprocessing choices, labels and exclusions. If those elements are undocumented, another team may be unable to reproduce the result or may apply the model to a population that was absent from the training material.

The same principle applies to public-sector and environmental datasets. A rainfall series might be adjusted after a station moves, a health dataset might be de-identified through several stages, or an agricultural dataset might change when sensors are recalibrated. Each decision can affect interpretation. Recording the decision is as important as retaining the final number.

Records Worth Preserving

  • Original files and a read-only preservation copy
  • Data dictionaries and codebooks linked to each release
  • Cleaning scripts, transformation logs and analytical code
  • Version dates, change notes and persistent identifiers
  • Information about missing data, exclusions, corrections and outliers

Provenance does not require every project to build an elaborate technical system. A well-structured README, a version-controlled repository and a change log may be enough for a modest study. Larger projects can use workflow tools, machine-readable metadata and automated capture. The standard should be proportional, but the basic principle should remain: users need to see how the released data came to be.

Documentation Carries Ethical Responsibilities

The conference’s focus on global food security and applied research also highlights a limit to purely technical documentation. A dataset can be findable and reproducible while still creating harm if its description ignores consent, power, privacy or cultural authority.

Australian research needs to take this seriously in work involving Aboriginal and Torres Strait Islander peoples. Documentation should identify who has authority over the data, what permissions were granted, whether community governance applies, and whether some materials should not be openly released. Indigenous data sovereignty is not solved by adding a generic ethics statement after the data has been collected.

Sensitive information can also arise in health, education, disability, migration and small-community research. A file may be de-identified yet remain identifiable when combined with other public sources. Documentation should explain the access model, disclosure controls, permitted uses and conditions for requesting restricted material. “Open” is not automatically the most responsible setting.

Australian research funders and institutions increasingly expect data management plans, retention decisions and evidence of responsible sharing. Researchers working under ARC or NHMRC expectations need documentation that reflects consent and governance from the beginning, rather than trying to reconstruct those details at publication time.

Questions For Ethical Data Records

  • Who created the data, and who has authority to describe or share it?
  • What consent, licence or community agreement governs reuse?
  • Could people, households, places or culturally significant information be identified?
  • Which uses are permitted, discouraged or prohibited?
  • What access pathway should a future researcher follow?
  • How should users acknowledge contributors, custodians and communities?

This approach fits local research realities. A dataset collected on Country is not simply a neutral collection of coordinates and observations. Its description may need to acknowledge Traditional Owners, community permissions and place-based knowledge. Clear documentation protects participants and helps future users understand that access carries responsibilities.

From Conference Insight To Daily Practice

The lasting value of IASSIST 2017 lies in connecting documentation with ordinary research work. A data record should begin when a project begins, not when a team is preparing a paper or trying to meet a repository deadline. Early notes are usually more accurate because collection decisions, unusual events and local terminology are still fresh.

A useful workflow starts with a data management plan. It identifies file formats, naming conventions, storage locations, backup arrangements, roles and likely users. During collection, researchers maintain a data dictionary and record deviations from the original protocol. During analysis, they preserve scripts, software versions and transformation decisions. Before sharing, they review disclosure risk, access terms, licences and the quality of the public description.

For a field team in regional Queensland, this may include recording when a sampling site became inaccessible after flooding, how local place names relate to map coordinates, and which instrument was used after a replacement. For a project based in Sydney or Adelaide, it may mean documenting survey recruitment, weighting decisions and changes to questionnaire wording. These details are easy to overlook and difficult to recover later.

Australian researchers also work across a mixed data market. Commercial platforms, university repositories, government portals and national services may all be involved. A team might use cloud storage during collection, an institutional repository for preservation, and a discipline archive for discovery. Documentation needs to remain portable so that the meaning of the data does not depend on one vendor or one staff member.

A Practical Documentation Routine

  • Create a project README before collecting the first observation
  • Agree on file names, formats, units, identifiers and version rules
  • Update the data dictionary whenever a variable or code changes
  • Record unusual events, decisions, exclusions and quality checks
  • Separate sensitive working files from material suitable for sharing
  • Deposit documentation beside the data and link it to related outputs

The conversational habits of Australian workplaces can make this routine effective when they are made explicit. A quick chat, a Teams message or an “arvo” debrief may solve an immediate problem, but it is not a durable record. The key decision should be transferred into the project documentation while the context is still clear.

IASSIST 2017 presented data as a common language. That language depends on definitions, grammar and shared conventions. Documentation supplies those elements, allowing data to travel between laboratories, repositories, communities and policy settings without becoming detached from its meaning. For Australian research, the lesson is straightforward: good records make data more discoverable, more accountable and far more useful over time.

At the Conference

What attendees experienced in Lawrence

p1 reed

Plenary Sessions

Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.

platinum icpsr

Workshops

Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.

gold rockhurst

Social Events

An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.

Venue & Accommodations

Where the conference took place

silver ifdo

Kansas Union

University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.

silver ciser

The Eldridge

701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.

platinum icpsr

Springhill & TownePlace Suites

Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.

Program Highlights

Sessions and activities

Plenary Sessions Workshops Poster Presentations Committee Meetings Opening Reception Banquet Tour Kansas City Pecha Kucha Check-In Local Favorites

Getting Here

Lawrence, Kansas

Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045

Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).

Plan Your Stay

Accommodation options that were available

silver ifdo

The Eldridge

701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.

silver ciser

Springhill Suites

Marriott property. Room block reserved under "KU IASSIST Conference."

gold rockhurst

TownePlace Suites

Marriott property. Room block reserved under "KU IASSIST Conference."