Abstract low-poly geometric background in shades of pale blue and white

Data in the Middle: The common language of research

Data Curation Challenges in Large-Scale Agricultural Studies

Modern agricultural science has become a data-intensive enterprise, with researchers routinely handling terabytes of sensor readings, satellite imagery, genomic sequences, and yield records collected across hundreds of trial sites. Nowhere is this complexity felt more sharply than in Australia, where institutions such as CSIRO, the University of Adelaide, and the Grains Research and Development Corporation coordinate field work across vast climatic gradients. From the black-soil plains of the Darling Downs to the irrigation districts of the Murray–Darling Basin, every paddock tells a different story, and stitching those stories together into a coherent dataset requires serious curation muscle.

The trouble is that agricultural data is messy in ways that laboratory datasets rarely are. Field trials run for decades under shifting protocols, sensors drift or get replaced mid-season, and growers often scribble notes on scraps of paper before transcribing them later over a cup of tea. Researchers who reckon their data will "she'll be right" frequently discover, years later, that critical provenance information has gone missing. Compounding the problem is the patchwork of funding bodies, universities, and private partners involved in any large Australian study, each with its own filing conventions and retention rules.

This piece walks through the main hurdles facing curation teams working on big agronomic projects and offers some field-tested approaches drawn from recent practice. It draws on the kinds of lessons surfaced during sessions at IASSIST 2017, where presenters wrestled openly with best practices for archiving conference materials, longitudinal datasets, and the awkward middle ground between active research and long-term preservation.

Why Agricultural Datasets Buckle Under Their Own Weight

Most research domains produce datasets that grow predictably. Agricultural studies do not. A single multi-site trial might combine soil chemistry from a lab in Perth, drone imagery collected by a contractor in Wagga Wagga, manually logged rainfall from a station near Longreach, and yield monitor data uploaded by a grower who swears the combine "works a treat" but never quite formats the files the same way twice. Multiply that across five seasons and a dozen sites and you have a dataset whose structure shifts subtly with every harvest.

The volume itself is part of the problem. High-resolution multispectral imagery from a single paddock can run into the hundreds of gigabytes, and time-series soil moisture probes generate millions of small files that defeat conventional folder-based organisation. Researchers at the Hawkesbury Institute for the Environment have publicly noted that simply locating a specific sensor reading from 2019 can take an afternoon when naming conventions drift. Curation teams therefore spend a surprising share of their time not on the data itself but on the metadata layer that lets anyone find anything in the first place.

Another weight comes from the sheer number of contributors. A typical GRDC-funded project might involve plant breeders, soil scientists, statisticians, extension officers, and half a dozen PhD students, all generating their own spreadsheets. Each person brings habits inherited from previous projects, and standardising those habits retroactively is harder than it sounds. The result is a curation backlog that grows faster than anyone can chase it.

Metadata Standards That Rarely Match

Agricultural research straddles several disciplinary metadata traditions, and choosing among them is rarely a clean decision. Geographic information communities favour ISO 19115, ecological datasets lean on the Ecological Metadata Language, agronomy trials often default to the Agricultural Metadata Element Set, and generic repositories tend to use Dublin Core or DataCite. The table below sketches how some of these schemas line up against the realities of a multi-site Australian trial.

Metadata field Dublin Core ISO 19115 AgMES DataCite
Spatial coverage Generic place name Bounding box, CRS Administrative region Free text
Temporal range date range date + duration cropping calendar publication year
Variables measured subject keywords feature catalogue crop, trait, protocol subject keywords
Provenance contributor lineage statement trial design block relatedIdentifier
Access conditions rights distribution info embargo flag rightsList
Suitable for sensor streams Limited Strong Moderate Limited

The mismatch shows up quickly in practice. A soil moisture time-series might be perfectly described under ISO 19115 but completely opaque when expressed only in Dublin Core. A plant breeding dataset fits AgMES well but chokes when it needs to describe drone-derived canopy height. Curation teams often end up implementing hybrid schemas, which works but introduces a maintenance burden that compounds over years.

Weather, Sensors, and the Trouble With Time

Australian agriculture runs on weather data, and weather data is unforgiving. Sensors fail in the heat of a January afternoon on the Liverpool Plains, loggers lose contact during a dust storm, and rain gauges get knocked over by curious livestock. A time-series that should be continuous ends up riddled with gaps that have to be flagged rather than imputed, because every gap carries meaning for agronomic interpretation. Researchers at the Tasmanian Institute of Agriculture have described weeks lost to reconciling logger clocks that drifted because nobody reset them at the end of daylight saving.

Continuous monitoring generates its own curation nightmare: file counts balloon. A single farm-scale weather station can write a file every five minutes for a decade, which adds up to more than a million tiny files. Many archival systems cope badly with this scale, and even modest repositories struggle to list, checksum, and back up such collections within a reasonable maintenance window. The honest fix is usually some form of aggregation at ingest, but that requires curation decisions made long before anyone knows which timestamps will matter.

Then there is the satellite and drone layer. Remotely sensed data arrives in georeferenced rasters that often need atmospheric correction, cloud masking, and tile reassembly before they can be linked to ground observations. Each processing step is a fork in the provenance chain, and every fork has to be documented or the entire dataset loses traceability. A grower might say the imagery "looks fair dinkum", but the research record needs something rather more rigorous.

Provenance Across Multiple Institutions

When a project involves CSIRO, a state department of primary industries, a university, and a handful of private collaborators, provenance becomes a governance problem as much as a technical one. Each partner holds data under its own ethics, funding, and commercial arrangements, and the act of merging records can breach any of those constraints. ABARES analysts working on national yield forecasts, for example, must respect the data-sharing rules of contributing farms even when those farms have long since been amalgamated into larger holdings.

Curation teams respond by building provenance graphs that explicitly link each dataset version back to its source files, processing scripts, and approval status. Tools like PROV-O and Research Objects handle the modelling side reasonably well, but the social side is harder. Researchers must agree on what counts as a meaningful version boundary, and that conversation tends to happen late, under deadline pressure, when somebody asks why the 2022 trial results differ from those in the internal report.

A common Australian workaround is to anchor provenance in the trial design document rather than the data files. If the design changes, a new design version is issued, and every downstream artefact references it. This makes the design document the canonical citation target and reduces the temptation to silently edit data files in response to reviewer comments. It also gives data managers a defensible answer when a colleague asks, months later, why a particular value is what it is.

Storage, Cost, and the Long View

Keeping agricultural data for the long term is expensive, and the cost falls unevenly. Universities in Sydney and Melbourne can lean on institutional repositories with mature preservation pipelines, while regional collaborators in places like Tamworth or Esperance often rely on a shared drive that nobody is quite sure gets backed up. The gap shows up in grant applications, where smaller partners quietly absorb storage costs that were never budgeted.

A pragmatic pattern emerging across Australian institutions is tiered storage: hot data sits on fast disk for the active research phase, warm data moves to object storage for the publication window, and cold data lands in a tape-backed archive for the long tail. The transitions need to be scripted and auditable, otherwise files get lost in the shuffle. Curation teams also increasingly use commercial cloud platforms for the hot tier, taking advantage of Australian data centres that have come online to satisfy data residency requirements.

Budget conversations remain difficult. Funders like the Australian Research Data Commons have done valuable work in supporting national infrastructure, but the per-project allocation rarely covers the decades-long horizon that good agricultural science actually demands. Researchers planning new trials are well advised to write a preservation budget into the original grant rather than discover the gap after the fieldwork wraps up.

FAIR Principles in Practice

The FAIR framework has been widely adopted in Australian agricultural research, but its operational meaning varies. Findable usually means a DOI minted through DataCite or ANDS, although some legacy datasets sit in institutional catalogues with no persistent identifier at all. Accessible tends to mean either open download or a request-mediated workflow, with embargo periods common for commercial partner data. Interoperability is where the gap is largest, because few projects have the resources to publish fully harmonised vocabularies. Reusable is the hardest, because reuse depends on rich contextual metadata that busy researchers rarely have time to write.

A useful local example comes from the dairy industry work coordinated through Dairy Australia, where shared trait dictionaries now let datasets from different herds be pooled for genetic analyses. The dictionaries were not glamorous to develop, but they unlocked analyses that simply could not have run before. The lesson generalises: the boring metadata work is what makes the science reproducible.

Implementation is usually incremental. A small curation team can lift a dataset from non-FAIR to partially FAIR by adding a single machine-readable licence file and a basic methods document, then build from there. Trying to do everything at once almost always stalls.

Workflows That Actually Work in the Field

Curation workflows tend to succeed when they are invisible to the busy researcher and unforgiving to the careless one. The patterns that have held up across Australian projects include automated ingest scripts that validate filenames and run basic checksums, a shared template for trial metadata that researchers fill in once and reuse, and quarterly curation sprints where the team reconciles outstanding provenance questions while the project is still fresh in everyone's mind.

Equally important is the social design. Naming a single data steward per trial, giving that person real authority over file formats, and protecting their time from competing demands tends to produce cleaner datasets than any technical fix. In the words of one New South Wales data manager, "if you want clean data, you have to fight for the boring bits," a sentiment that captures the reality of large-scale curation work. Tools help, but the discipline is human.

Recommendations for Curation Teams Planning Large Studies

  • Negotiate a shared metadata template before the first plot is sown, and bind it to the data management plan from day one.
  • Assign a named data steward with real time allocated in the grant, not as a courtesy title.
  • Adopt tiered storage with scripted transitions, and document the policy in a place researchers will actually read.
  • Anchor provenance in the trial design document so that version boundaries align with scientific decisions.
  • Reserve a slice of the budget specifically for long-term preservation, calculated across the realistic lifetime of the dataset.
  • Publish a machine-readable licence with every public release to remove ambiguity for downstream users.
  • Run quarterly reconciliation sprints while memory of the fieldwork is still fresh, rather than during the manuscript rush.

At the Conference

What attendees experienced in Lawrence

p1 reed

Plenary Sessions

Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.

platinum icpsr

Workshops

Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.

gold rockhurst

Social Events

An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.

Venue & Accommodations

Where the conference took place

silver ifdo

Kansas Union

University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.

silver ciser

The Eldridge

701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.

platinum icpsr

Springhill & TownePlace Suites

Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.

Program Highlights

Sessions and activities

Plenary Sessions Workshops Poster Presentations Committee Meetings Opening Reception Banquet Tour Kansas City Pecha Kucha Check-In Local Favorites

Getting Here

Lawrence, Kansas

Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045

Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).

Plan Your Stay

Accommodation options that were available

silver ifdo

The Eldridge

701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.

silver ciser

Springhill Suites

Marriott property. Room block reserved under "KU IASSIST Conference."

gold rockhurst

TownePlace Suites

Marriott property. Room block reserved under "KU IASSIST Conference."