Abstract low-poly geometric background in shades of pale blue and white

Data in the Middle: The common language of research

Navigating Agricultural Data From Field to Farm Decision

Agricultural research produces data in many forms, from soil moisture readings and satellite imagery to livestock records, crop trials, weather observations and interviews with growers. Each source carries its own assumptions, timing, scale and level of uncertainty. Treating these materials as a connected evidence stream helps researchers preserve meaning as information moves from collection to analysis and practical use.

The data lifecycle begins before a sensor is switched on. Research teams need to decide what they are measuring, why the information is needed, who may access it, and how results will be maintained after a project ends. These choices influence sampling design, metadata, consent, storage costs and the credibility of later findings.

For Australian agriculture, the setting is especially varied. A project may combine observations from irrigated horticulture near Mildura, broadacre cropping in regional New South Wales, cattle stations in Queensland and remote farms outside Perth. The distance between sites, uneven connectivity, drought cycles and different production systems make standardisation valuable, while local knowledge keeps standardisation from becoming simplistic.

The archived IASSIST 2017 programme provides a useful intellectual backdrop for this work. Its focus on data as a common language connects agricultural science with librarianship, computing, social research and policy. That shared vocabulary is essential when a plant scientist, statistician, farmer, software engineer and government analyst need to interpret the same dataset responsibly.

Defining Purpose, Scope and Stewardship

A sound agricultural data management plan starts with a research question that can guide collection. “Improve water efficiency” is too broad to determine which variables belong in a dataset. A more useful objective might examine how irrigation scheduling affects yield and soil salinity across specified soil types and seasons. The narrower question identifies the observations, comparison groups and time period required.

Researchers should map stakeholders at this stage. Growers may provide operational records, agronomists may interpret field conditions, Indigenous communities may hold place-based knowledge, and public agencies may supply climate or land-use data. Each contributor can have different expectations about attribution, access and future reuse. Agreements should record these expectations before data begins to accumulate.

Scope also includes practical constraints. A sensor network that generates readings every minute may exceed the project’s storage or quality-control capacity, while a survey designed for large commercial farms may exclude smaller family operations. Clear decisions about resolution, sampling frequency and geographic coverage prevent researchers from collecting impressive volumes of information with little analytical value.

Australian projects must consider legal and ethical responsibilities alongside scientific ones. The Privacy Act 1988 may apply when farm employees, landholders or identifiable businesses can be linked to records. Biosecurity requirements can affect the movement of samples, equipment and biological material across borders or between jurisdictions. Indigenous data governance principles may require community authority over collection, interpretation and sharing.

Capturing Reliable Evidence in the Field

Field data is shaped by its collection environment. Heat, dust, poor mobile coverage, inconsistent power and long travel distances can disrupt devices and workflows. A vineyard near Adelaide may support frequent uploads, while a grazing property in the Northern Territory may require local storage and periodic transfer. Designing for offline operation, battery limits and manual backup is often more realistic than assuming continuous connectivity.

Metadata gives measurements their meaning. A soil moisture value needs a timestamp, depth, coordinate, instrument type, calibration status and unit. A yield figure needs information about the harvested area, moisture content, cultivar and weighing method. Without these details, datasets that appear compatible may produce misleading comparisons.

Field teams benefit from documented protocols and controlled vocabularies. Terms such as “rainfall,” “effective rainfall,” “irrigation event” and “water application” should have defined meanings. Versioned forms, standard file names and validation rules help identify missing or implausible values while the collection process is still active.

Practical Controls for Field Collection

  • Record location, time zone, unit, instrument and operator with each observation.
  • Use calibration schedules and maintain an audit trail for repairs or replacements.
  • Keep a local copy when internet access is unreliable, then reconcile records after synchronisation.
  • Test forms and sensor workflows at representative sites before scaling up.
  • Mark estimated, imputed, below-detection and missing values distinctly.

Quality assurance should be proportionate to the research purpose. A national crop forecast may require automated checks across millions of observations, whereas a small agronomy trial may depend on careful field notes and duplicate measurements. Either way, researchers should preserve the original record and document every correction rather than silently overwriting it.

Storing, Linking and Protecting Agricultural Information

Agricultural datasets commonly sit across spreadsheets, laboratory systems, cloud platforms, farm-management software and government repositories. A lifecycle approach creates an inventory showing where each asset is held, who owns it, how often it changes and which other data it depends on. This inventory becomes particularly important when a project ends and staff, contracts or software systems change.

File formats and identifiers determine whether data can be reused. Open, well-documented formats are generally preferable for long-term preservation, although proprietary systems may remain necessary during active farm operations. Persistent identifiers for sites, trials, samples and instruments help link records without relying on fragile names such as “North paddock final revised.”

Access controls should match sensitivity. Public weather observations can usually be shared widely, but farm financial records, personally identifiable information, commercial yields and geospatial details may need restricted access. De-identification is useful, though it does not guarantee anonymity when a small property can be recognised through location, crop type or production scale.

Australian researchers should also examine where cloud services store information and which provider can access it. Institutional policies, grant conditions, the Privacy Act and contractual obligations may affect cross-border transfers. A data management plan should state retention periods, backup arrangements, recovery procedures and the conditions for secure deletion.

Signals of a Healthy Data Asset

  • A new researcher can understand the dataset without relying on its original creator.
  • Definitions, units, transformations and quality flags are documented.
  • Raw data is protected while cleaned and derived versions remain traceable.
  • Permissions and consent conditions are visible to authorised users.
  • Backups are tested rather than assumed to work.

Data linkage adds value when it is governed carefully. Combining paddock boundaries with satellite imagery, commodity prices and Bureau of Meteorology observations can reveal patterns that no single source provides. It can also create new risks, including re-identification, inappropriate inference or conclusions that ignore differences in timing and geographic scale.

Analysing Data With Context and Care

Analysis should begin with an understanding of how the data was produced. Missing rainfall observations may reflect an instrument failure, a site being inaccessible after flooding or a deliberate change in sampling. Each explanation has different implications for modelling. Treating all missing values as the same can distort estimates of crop performance or climate impact.

Researchers working with agricultural data often combine statistical modelling, geospatial analysis, machine learning and qualitative interpretation. Deep learning may classify satellite imagery or detect disease symptoms, but its predictions still depend on representative training data and transparent validation. A model trained largely on Victorian wheat fields may perform poorly on Western Australian varieties, tropical crops or unusual seasons.

Uncertainty should be communicated in terms that decision-makers can use. Farmers may need a prediction interval, a map of confidence, or a clear warning that a recommendation applies only to certain soil and climate conditions. Publishing a single precise number can create false confidence when the underlying evidence is variable.

Australian context matters in interpretation. A dry season in the Murray–Darling Basin, a cyclone-affected period in northern Queensland and a heatwave around Perth create different pressures on farms and datasets. Market conditions also shape outcomes: a profitable response to a premium cherry market may not translate to a grain enterprise supplying export contracts. Analysis should distinguish environmental effects from prices, labour, infrastructure and management choices.

Reproducible workflows make this reasoning inspectable. Scripts, model versions, parameter settings and intermediate datasets should be retained alongside research notes. When an algorithm is updated, the team should be able to identify which results changed and why. This protects scientific integrity and makes collaboration easier across universities, government departments and industry organisations.

Sharing Results and Turning Evidence Into Action

Data becomes useful when its findings can travel beyond the original project without losing essential context. Researchers should prepare documentation for different audiences: a technical data dictionary for analysts, a plain-language summary for growers, a methods note for peer reviewers and a policy brief for public agencies. These outputs may use the same evidence while answering different practical needs.

Repositories and institutional archives support discovery, preservation and citation. A conference archive such as conference resources demonstrates how programme materials, presentations and practical information can remain useful after an event has finished. Agricultural projects can apply a similar principle by preserving datasets with stable documentation rather than leaving important knowledge in personal drives or short-lived project websites.

Sharing should respect the original agreement with contributors. Open data is valuable for public research, yet controlled access may be appropriate for commercially sensitive farm records, threatened species locations or information that could expose a vulnerable community. Licences should explain what users may do with the material, and acknowledgements should recognise growers, technicians, data custodians and community partners.

A useful release includes provenance: where the data came from, how it was cleaned, what limitations apply and which version supports the published result. This helps users judge whether evidence suits a new crop, region or season. In Australia, a recommendation developed for irrigated production near Melbourne may require substantial adaptation before use in a remote dryland system.

Comparing Data Lifecycle Priorities

Lifecycle stage Main question Agricultural example Evidence of good practice
Planning What decision or research question will the data support? Testing irrigation strategies under different soil conditions Defined variables, scope, stakeholders and consent
Collection Are observations consistent and sufficiently documented? Recording soil moisture, crop stage and irrigation events Calibrated instruments, protocols and field metadata
Management Can information be protected, found and reused? Linking sensor files with trial and paddock identifiers Version control, backups, access rules and clear formats
Analysis Are results valid across sites, seasons and production systems? Modelling yield response during variable rainfall Quality checks, uncertainty reporting and reproducible code
Sharing Can others interpret and apply the findings responsibly? Publishing a dataset with a grower-facing summary Provenance, licence, limitations and appropriate access

The lifecycle is iterative rather than a straight line. Analysis may reveal that a variable was collected at the wrong scale, while farmer feedback may show that a technically accurate dashboard is too slow or complex for seasonal work. Those findings should lead to revised collection protocols, improved metadata and better communication in the next cycle.

Successful agricultural research therefore depends on relationships as much as infrastructure. A common language allows specialists to negotiate definitions, recognise uncertainty and preserve the experiences behind numerical records. When data is planned, captured, protected, analysed and shared as a connected process, research findings have a stronger chance of supporting resilient farms, accountable policy and practical decisions across Australia.

At the Conference

What attendees experienced in Lawrence

p1 reed

Plenary Sessions

Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.

platinum icpsr

Workshops

Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.

gold rockhurst

Social Events

An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.

Venue & Accommodations

Where the conference took place

silver ifdo

Kansas Union

University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.

silver ciser

The Eldridge

701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.

platinum icpsr

Springhill & TownePlace Suites

Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.

Program Highlights

Sessions and activities

Plenary Sessions Workshops Poster Presentations Committee Meetings Opening Reception Banquet Tour Kansas City Pecha Kucha Check-In Local Favorites

Getting Here

Lawrence, Kansas

Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045

Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).

Plan Your Stay

Accommodation options that were available

silver ifdo

The Eldridge

701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.

silver ciser

Springhill Suites

Marriott property. Room block reserved under "KU IASSIST Conference."

gold rockhurst

TownePlace Suites

Marriott property. Room block reserved under "KU IASSIST Conference."