Reproducibility in agricultural data: standards shaping practice
Agricultural science has long relied on field trials, long-term monitoring plots, and observational networks that span decades. Reproducibility in this field is uniquely difficult because experiments depend on soil, climate, and seasonal cycles that cannot be replayed. A wheat trial sown in Horsham during 2014 exists only in the records, samples, and notebooks that survive from that season. When researchers revisit the same plots years later, they rely entirely on the quality of stored data to compare outcomes, interpret trends, and refine models. Reproducibility is therefore a foundational requirement rather than a methodological preference; it is the substrate on which multi-decadal agronomic knowledge is built.
The 2017 IASSIST conference in Lawrence, Kansas, gathered data professionals, librarians, and domain specialists under the theme "Data in the Middle: The Common Language of Research." Agricultural data featured prominently because the sector sits at the intersection of environmental variability, global food security, and rapidly evolving analytical methods. Presenters emphasised that reproducibility in agriculture requires more than open files. It demands shared vocabularies, machine-readable metadata, transparent workflows, and durable identifiers that survive software obsolescence and institutional change.
Australian researchers bring a particular sensibility to this conversation. Work conducted through organisations such as the CSIRO, the Grains Research and Development Corporation, and the state departments of primary industries combines paddock-scale experimentation with continental-scale environmental monitoring. Projects such as TERN, the Terrestrial Ecosystem Research Network, and the Long Paddock drought service generate terabytes of time-series data that must remain interpretable across decades. Reproducibility, in the Australian context, means ensuring that a yield comparison made today can still be validated by a graduate student in 2040 using archived datasets and the documentation that accompanied them.
The standards emerging from the IASSIST discussions offer practical guidance for anyone working with agricultural data. They address metadata structure, identifier systems, workflow capture, and repository governance. What follows is a synthesis of those discussions, framed for Australian practitioners balancing local research realities with the expectations of international collaborators and funders.
Reproducibility frameworks for agricultural research
The FAIR principles — Findable, Accessible, Interoperable, and Reusable — provide the most widely accepted framework for reproducible research data. Each principle translates into concrete actions: assigning persistent identifiers such as DOIs, using standardised metadata, depositing data in trusted repositories, and licensing outputs in ways that allow reuse. Agricultural datasets frequently fail the FAIR test because metadata is captured informally, files are stored on personal drives, and field protocols are documented only in spreadsheets that disappear when a project ends.
Several frameworks extend FAIR with agricultural specifics. The GODAN Standards Toolkit, the CGIAR open data guidelines, and the Research Data Alliance agricultural interest group have each proposed layered recommendations. These frameworks share a recognition that agricultural metadata must capture location, soil type, cultivar, sowing date, fertiliser regime, weather conditions, and management practices. Without these contextual variables, a yield number is meaningless regardless of how cleanly the dataset itself is formatted.
| Standard or toolkit | Primary focus | Typical users | Relevance to agricultural reproducibility |
|---|---|---|---|
| AGROVOC | Multilingual controlled vocabulary | Librarians, data managers | Aligns crop, soil, and livestock terminology across regions |
| GACS | Global Agricultural Concept Scheme | Research networks, repositories | Bridges AGROVOC, NAL Thesaurus, and CAB Thesaurus |
| DataCite | Persistent identifier and metadata schema | Repositories, publishers | Issues DOIs so datasets can be cited and tracked over time |
| DCAT | Dataset catalogue metadata model | Government agencies, data portals | Enables cross-portal harvesting of dataset descriptions |
| ISO 19115 | Geographic information metadata | Spatial analysts, environmental agencies | Documents location, projection, and sampling geometry |
| Crop Ontology | Trait and phenotype vocabulary | Plant breeders, phenomics platforms | Standardises trait names for genotypic and phenotypic studies |
The table is not a ranking. It illustrates how different tools occupy different layers of the reproducibility stack. A single Australian trial might rely on the Crop Ontology for trait names, ISO 19115 for spatial context, and DataCite for citation, working together to make the underlying data reusable by researchers who were not present at sowing.
Standards presented at the Lawrence gathering
Sessions in Lawrence drew attention to standards that had moved from discussion documents to operational tools. The FAIRsharing registry, the re3data repository directory, and the FORCE11 software citation principles were presented as ready-to-implement resources. Speakers demonstrated how journals and funders were beginning to require data availability statements aligned with these standards, shifting reproducibility from an academic ideal to an administrative expectation.
A recurring theme was the tension between flexibility and standardisation. Agricultural research is highly local: a vineyard trial in the Barossa Valley and a cotton experiment near Narrabri differ in almost every variable except the underlying need for comparable documentation. Standards must be loose enough to accommodate local reality and strict enough to permit automated aggregation. The conference highlighted vocabularies such as the Environment Ontology and the Plant Ontology as examples of community-maintained resources that achieve this balance.
Speakers also stressed the role of social infrastructure. Standards succeed when a community of practice supports them: trainers who run workshops, curators who maintain term lists, and reviewers who recognise compliant submissions. Without that social scaffolding, even well-designed schemas fall out of use as researchers revert to whatever is easiest in the moment.
Metadata schemas and controlled vocabularies
Metadata is the connective tissue of reproducible research. For agricultural data, the schema must record who collected what, where, when, how, and under what conditions. Generic schemas such as Dublin Core cover the basics, but domain-specific extensions are essential. The Australian National Data Service, now operating within the Australian Research Data Commons, has promoted profiles built on DataCite and DCAT, adding fields for site identifiers, treatment codes, and instrument calibration. These profiles are often invisible to end users but crucial for anyone attempting to merge datasets collected by different research groups.
Controlled vocabularies are equally important. A "drought" recorded by a NSW researcher may not match what a colleague in Western Australia considers drought, since rainfall baselines differ substantially between regions. Vocabularies like GACS attempt to harmonise terminology across languages and traditions, but local curation remains necessary. Australian institutions have invested in vocabularies for native flora, soil classifications, and pest species that complement international resources. The practical lesson from the IASSIST discussions was that vocabulary work is never finished; it requires ongoing governance as new crops, sensors, and analytical methods appear.
Metadata quality is also a matter of timing. Capturing protocol details at the moment of sowing, or instrument serial numbers when a probe is deployed, is far more reliable than reconstructing them years later from memory. Data management plans work best when treated as living documents that travel with the project rather than compliance artefacts filed at grant submission.
Computational workflows and shared pipelines
Even the cleanest dataset is not reproducible if the analysis that produced the published figure cannot be reconstructed. Computational reproducibility depends on capturing the software environment, the sequence of operations, and the input parameters used. Tools such as R Markdown, Jupyter notebooks, Snakemake, and Nextflow were discussed at Lawrence as practical means of turning ad hoc analyses into documented pipelines. For Australian agricultural researchers, these tools intersect with platforms such as the APSIM crop simulation model and high-performance computing through the National Computational Infrastructure.
A workflow captured in a version-controlled repository, executed through a workflow manager, and published alongside the data satisfies most reproducibility expectations. The conference emphasised that adoption of these tools is often cultural rather than technical. Senior researchers who learned analysis through spreadsheets may resist moving to notebooks, while early-career scientists who grew up with Git may underestimate the documentation needs of field scientists. Bridging this gap requires institutional support, training, and recognition that workflow work counts as research output.
Containerisation and environment capture were discussed as the next frontier. Bundling an analysis with its software dependencies in a container image guarantees that the same code runs the same way on different machines, offering a form of digital preservation that complements archival metadata for long-term agricultural studies.
Local adoption across Australian institutions
Australia has built a distinctive ecosystem for research data management. ANDS, now part of the Australian Research Data Commons, has worked for more than a decade to make FAIR principles practical for universities and government agencies. The ARDC data and compute platforms, alongside NCRIS-funded infrastructure, give Australian researchers access to repositories, identifier services, and training programs that smaller national systems struggle to provide. Agricultural users benefit from domain-specific investments such as the Australian Gridded Climate Data product, the Soil and Landscape Grid of Australia, and supply-chain datasets increasingly shared with research partners.
At the institutional level, the University of Sydney, the University of Queensland, the University of Adelaide, La Trobe University, and others have established data stewardship teams that work directly with agricultural faculties. These teams help researchers deposit data, write metadata, and select licences. Their work is often invisible in the published literature but critical to the reproducibility of any long-running trial, whether it tracks wheat rust resistance in the Wimmera or pasture productivity in the New England Tablelands.
State agencies contribute their own layers. Agriculture Victoria, the NSW Department of Primary Industries, and the Queensland Department of Agriculture and Fisheries publish datasets through portals that increasingly align with federal standards. National coordination through the ARDC and the Australian Bureau of Agricultural and Resource Economics and Sciences has begun to address the remaining gap between state and university-held datasets, although progress remains uneven.
Pathways for sustainable reproducibility
Standards alone do not produce reproducibility; they require sustained practice. The IASSIST conversations pointed to several levers that move standards from paper to everyday use. Funders that require data management plans and reward their execution create accountability. Journals that require data availability statements and enforce them build reviewer habits that travel into the lab. Institutions that count data papers, software citations, and metadata curation as legitimate scholarly outputs shift career incentives. Each lever reinforces the others.
For Australian agricultural research, sustainability also depends on partnership with industry. Grower groups such as Grain Producers Australia, the National Farmers' Federation, and regional bodies like Birchip Cropping Group generate valuable observational data that only becomes reproducible when collected against agreed schemas. The cotton industry's CottonMap, the dairy industry's DataGene, and the sugar industry's Yield Decline Joint Venture are early examples of industry-academic collaboration on data standards. Extending this pattern more widely would lift reproducibility across the sector without imposing burdens on any single research group.
The standards discussed at IASSIST 2017 remain relevant because the underlying problems have not changed. Field trials still run on seasonal calendars, sensors still produce streams of numbers, and models still depend on inputs that must be traceable. What has changed is the willingness of the research community to treat data work as a first-class scholarly activity. Australian researchers who adopt this view, supported by infrastructure that already exists, can deliver agricultural science that is genuinely reproducible — and therefore genuinely trustworthy.
Practical steps for Australian research teams
- Deposit datasets in a domain-appropriate repository such as the CSIRO Data Access Portal, an ARDC-supported platform, or an international repository like Dryad, and mint DOIs through DataCite.
- Adopt a metadata profile that extends DataCite or DCAT with agricultural fields for location, soil type, cultivar, treatment, and protocol version.
- Use controlled vocabularies such as AGROVOC, GACS, and the Crop Ontology for traits, and pair them with the Australian Soil Classification for soil descriptors.
- Capture analyses in version-controlled scripts or notebooks, and pair them with environment specifications using container images or package locks.
- Include a data availability statement in every publication, citing the dataset DOI and the licence under which reuse is permitted.
- Coordinate with institutional data stewards or ARDC data champions when planning a project, so that metadata work begins at trial design rather than at publication time.
- Engage with industry partners and grower groups early, agreeing on field-data schemas before the first paddock is sown so that industry-academic datasets are interoperable from day one.
At the Conference
What attendees experienced in Lawrence
Plenary Sessions
Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.
Workshops
Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.
Social Events
An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.
Venue & Accommodations
Where the conference took place
Kansas Union
University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.
The Oread
1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.
The Eldridge
701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.
Springhill & TownePlace Suites
Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.
Program Highlights
Sessions and activities
Getting Here
Lawrence, Kansas
Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045
Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).
Plan Your Stay
Accommodation options that were available
The Eldridge
701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.
The Oread
1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.
Springhill Suites
Marriott property. Room block reserved under "KU IASSIST Conference."
TownePlace Suites
Marriott property. Room block reserved under "KU IASSIST Conference."