Abstract low-poly geometric background in shades of pale blue and white

Data in the Middle: The common language of research

Drafting a Data Management Plan for Digital Agriculture Projects

Digital agriculture is no longer a fringe curiosity on Australian research agendas. From autonomous weed-robots trialled at the University of Sydney's Narrabri site to satellite pasture biomass mapping used by graziers in western Queensland, the volume and variety of data flowing out of the paddock has grown beyond anything a paper notebook could contain. A well-formed data management plan gives that data a predictable pathway from the sensor to the archive and, just as importantly, gives funders, collaborators and future re-users a clear statement of intent.

The same plan reduces friction inside a project. When agronomists, software engineers and farm staff scattered between regional Victoria and the Riverina all know where raw files live, who owns the cleaned datasets and which embargo rules apply, meetings shift from clarifying basics to making decisions. A data management plan is therefore both a compliance artefact for institutions such as the Australian Research Council and a working document for everyday operations.

Australia's agricultural research sector stretches across tropical cane fields near Cairns, dryland wheat operations on the Eyre Peninsula and temperate dairy farms south of Melbourne. Each region brings its own data flavours: high-frequency soil moisture logs, drone-derived multispectral imagery, RFID tag reads from saleyards, and chemical residue test results. A single project may pull from several of these streams, which makes the planning step especially valuable before the first byte is collected.

The sections below walk through the choices a team is likely to face, from scoping the dataset inventory to selecting a long-term home for the outputs. The aim is a plan that reads well to a reviewer in Canberra and still makes sense to a research officer standing next to a weather station at dusk in Tamworth.

Scoping the Project and Listing the Data Assets

Every useful plan starts with a clear-eyed inventory. Rather than guessing what data a project might produce, walk through each work package and ask what files, tables, images or streams the team expects to handle. A soil moisture study on a property near Wagga Wagga, for example, may generate raw logger dumps, cleaned CSV tables, calibration certificates for each probe, and the publication-ready figures that flow from the analysis. Listing these artefacts early prevents the classic late-project scramble when a thesis student graduates and the only copy of an intermediate dataset sits on a departing laptop.

Project scoping also clarifies what the data is not. A grazing trial might collect biomass cuts but explicitly exclude drone footage because of budget. Writing those exclusions into the plan signals discipline to reviewers and protects the team from being asked later why certain materials are missing from the deposit.

A simple three-column table is often enough at this stage: data type, expected volume, and the instrument or person generating it. The table is not a deliverable in itself but it becomes the spine of the metadata section and the storage-cost estimate. Keep it short, revisit it monthly, and prune anything that turned out to be a dead end.

Following the Data Through Its Lifecycle

Digital agriculture data moves through recognisable phases: acquisition, validation, analysis, publication and preservation. The plan should describe each phase in plain language and name the tools, file formats and quality checks applied at each step. Acquisition from a LoRaWAN-connected rain gauge near Bundaberg, for instance, may produce raw binary packets that need decoding before they become readable weather records. Naming the decoder script and the version of the firmware in the plan means a colleague two years later can reproduce the conversion.

Validation deserves particular attention because agricultural data is notoriously noisy. A herd-weight sensor can produce biologically implausible values when an animal pushes against a gate, and a satellite-derived normalised difference vegetation index will misbehave at cloud edges. The plan should state how outliers are flagged, who reviews the flagging, and whether the original raw values are retained alongside the cleaned dataset. Retaining raw inputs aligns with the FAIR principle of keeping data findable, accessible, interoperable and reusable.

Analysis and publication phases benefit from explicit version control statements. A study using machine learning to forecast sorghum yield near Dalby might describe which model artefacts are kept, which training data subsets are released, and which hyperparameter logs accompany the model. Without these statements, the published paper becomes an orphan, hard to reproduce or extend.

Choosing Metadata Standards That Travel

Metadata is the contract between the dataset and anyone who encounters it later. In Australia, two reference points are useful. The first is the Australian National Data Service (ANDS) vocabulary, which provides practical guidance on descriptive fields and persistent identifiers. The second is the AgTrials and CGIAR-derived agronomic metadata schemas, which handle field-experiment variables like crop, cultivar, planting date and fertiliser regime. Many teams blend both, using ANDS for discovery metadata and the agronomic schema for the dataset body.

Standardisation should extend to units. Australian farming literature still mixes metric and imperial in older reports, and chemical concentrations can appear in parts per million, milligrams per kilogram or percent depending on the lab. The plan should mandate a single unit system for the project, name the controlled vocabularies used, and reference a date format such as ISO 8601. These small decisions remove hours of confusion during data integration.

File format choices matter too. Preferring open, well-documented formats such as NetCDF, CSV with a documented header, GeoTIFF and PDF/A for reports means the data can be opened by future researchers without proprietary licences. Where proprietary formats are unavoidable, the plan should explain why and describe a migration pathway.

Common tools for building a digital agriculture data management plan

Tool or service Best fit Strengths Watch out for
DMPTool (with Australian institutional templates) University-led projects Step-by-step prompts, funder templates Generic prompts need local tailoring
ARDC Data Management Planning guide Mixed teams new to planning Plain-English examples, Australian focus Less automation than dedicated tools
Open Science Framework Cross-institution collaboration Built-in storage, versioning, DOI minting Per-project storage quotas
RSpace Labs tied to electronic lab notebooks Tight integration with experiment records Subscription cost for large teams
CKAN-based institutional repositories Public-good datasets Strong discovery, open licences Requires metadata curation effort

Controlled vocabularies worth adopting

  • Australian Soil Classification for soil orders, depths and textures.
  • ANZSIC 2006 for industry and farm enterprise coding.
  • Australian Plant Census names for cultivated and native species.
  • ISO 8601 date-time formatting across all field records.
  • Creative Commons Australia licences matched to the intended reuse.

Storage, Backup and Sovereignty Considerations

Storage planning in Australia must answer three questions: where the active copy sits, where the backup lives, and whether either copy crosses an international border. Many universities already store active data on institutional infrastructure, but cloud platforms such as AWS Sydney, Azure Australia Central or local providers like Vault Systems offer alternatives with data residency guarantees. Whatever the choice, the plan should record the provider, the region and the retention period.

Backups are often treated as an afterthought in field-heavy projects. A reliable pattern is the three-two-one rule: three copies, on two different media, with one off-site. For a broadacre project collecting imagery over several seasons across the Wimmera, this might mean an on-farm network-attached storage array for immediate use, an institutional research data store for weekly sync, and an off-site cloud archive for disaster recovery. The plan should name each tier and the cadence of transfer.

Data sovereignty is a live topic. The Privacy Act 1988 and the Notifiable Data Breaches scheme impose obligations on any handling of personal information, and agricultural projects that include farmer names, addresses or financial records fall within scope. Even purely environmental projects should note where the data resides, because partners and funders may have contractual residency clauses. Naming the storage region explicitly protects the project from accidental non-compliance.

Sharing, Ethics and the Privacy Act

Sharing agricultural data unlocks the kind of cross-region comparisons that single trials cannot deliver. The CSIRO, state departments and the Australian Institute of Marine Science have shown the value of pooled datasets for rangelands, soil carbon and water-use efficiency. A good plan names the repository where the data will land, the licence under which it will be released (often Creative Commons Attribution for derivative-friendly sharing or CC-BY-NC where commercial sensitivities exist), and the embargo window during which the project team has privileged access.

Ethics review intersects with data planning whenever human participants are involved. Survey responses from landholders in the Murrumbidgee catchment, interviews with agronomists, or focus groups with Indigenous co-researchers all require Human Research Ethics Committee approval. The plan should reference the ethics approval number, the consent forms used, and the de-identification strategy. Where Indigenous data sovereignty principles apply, the CARE framework (Collective benefit, Authority to control, Responsibility, Ethics) should be cited and the plan should describe how Traditional Owners will be involved in decisions about reuse.

Sensitive commercial data, such as yield maps from a contract grower or proprietary seed trial results, also needs handling rules. The plan can list commercial-in-confidence categories and describe how they will be aggregated, masked or withheld before any public release. Clarity here prevents awkward late-stage negotiations with industry partners.

Practical steps for the first dataset deposit

  • Mint a persistent identifier such as a DOI through the chosen repository before the paper is submitted.
  • Upload a README alongside each dataset explaining column headings, units and known limitations.
  • Link the deposit record from the related publication and from the project's own page.
  • Notify co-authors and industry partners once the embargo lifts.
  • Add the deposit details to the final version of the data management plan.

Roles, Responsibilities and Governance

A data management plan that names no one will drift. The plan should identify a data steward for the project, a backup steward, and clear lines for each major decision: who approves metadata changes, who can release a dataset, who handles a breach notification. In small teams this may be the same person wearing several hats, but writing the role out loud makes the workload visible.

Distributed teams, common in multi-institution projects linking a Brisbane-based modeller with a Perth field team and a Hobart statistician, benefit from a short governance section. It can describe the cadence of data review meetings, the communication channel for urgent issues, and the escalation path if a partner institution changes its policies. Even a single paragraph prevents the all-too-common scenario where a dataset is uploaded under a personal account and becomes inaccessible when that staff member leaves.

Training is part of governance too. Listing induction steps for new team members, the location of the data dictionary, and the expected file-naming conventions keeps the project legible to rotating honours students and visiting researchers. A short onboarding checklist attached to the plan often does more for data quality than any technical control.

Reviewing, Archiving and Learning From the Plan

Plans decay quickly when treated as one-off grant attachments. Build a review cycle into the project, perhaps quarterly or aligned with each field season, and record changes in a version history table at the front of the document. Reviewers and auditors appreciate seeing how the plan evolved with the work.

Archiving is the final test of any plan. Before the project closes, the team should deposit datasets in a trusted repository such as the Australian Research Data Commons supported services, a domain-specific archive like the Terrestrial Ecosystem Research Network for environmental data, or an international partner for globally relevant datasets. The deposit record, including a persistent identifier such as a DOI, should be added to the plan and to any related publications.

The last step, often skipped, is a short lessons-learned note. What worked, what tripped the team up, and what would be done differently next time. Capturing that knowledge feeds directly into the next data management plan and, over time, builds a library of practice that lifts the whole Australian digital agriculture research community.

At the Conference

What attendees experienced in Lawrence

p1 reed

Plenary Sessions

Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.

platinum icpsr

Workshops

Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.

gold rockhurst

Social Events

An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.

Venue & Accommodations

Where the conference took place

silver ifdo

Kansas Union

University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.

silver ciser

The Eldridge

701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.

platinum icpsr

Springhill & TownePlace Suites

Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.

Program Highlights

Sessions and activities

Plenary Sessions Workshops Poster Presentations Committee Meetings Opening Reception Banquet Tour Kansas City Pecha Kucha Check-In Local Favorites

Getting Here

Lawrence, Kansas

Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045

Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).

Plan Your Stay

Accommodation options that were available

silver ifdo

The Eldridge

701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.

silver ciser

Springhill Suites

Marriott property. Room block reserved under "KU IASSIST Conference."

gold rockhurst

TownePlace Suites

Marriott property. Room block reserved under "KU IASSIST Conference."