Abstract low-poly geometric background in shades of pale blue and white

Data in the Middle: The common language of research

Data standards and the digital farm after IASSIST 2017

The IASSIST 2017 conference, held at the University of Kansas in Lawrence from 23 to 26 May, explored how data can become a common language across research communities. Its archived programme connected familiar library and information-management concerns with emerging fields such as big data, deep learning, digital agriculture and global food security. That combination made agriculture a useful test case for a wider question: how can data produced by different people, machines and institutions be understood and reused together?

Digital agriculture depends on this shared meaning. A soil sensor, a satellite platform, a farm management system and a government statistics agency may all describe the same paddock, crop or season in different ways. Standards provide the bridge between those systems. They help identify what was measured, where and when it was recorded, which units were used, who can access it and whether the result can be trusted.

Data concern Practical agricultural example Standardisation need
Meaning “Yield” may refer to harvested weight, estimated yield or yield per hectare Shared definitions and controlled vocabularies
Location A farm may be represented by an address, boundary, GPS point or cadastral parcel Geospatial metadata and persistent identifiers
Time A crop observation may use local time, UTC, a growing season or a harvest year Consistent date formats and temporal descriptions
Measurement Moisture, rainfall and fertiliser rates may use different units Machine-readable units, calibration details and provenance
Access Commercial, public and Indigenous data may have different rights Clear licences, permissions and governance rules
Reuse A dataset may be copied into a model or combined with satellite imagery APIs, interoperable formats and documented lineage

Why agricultural data needs a common language

A modern agricultural operation generates data through many channels. Machinery records seeding rates and fuel use, weather stations capture rainfall and temperature, satellites identify crop stress, laboratories analyse soil and mobile applications track livestock or chemical applications. Each source has value, yet value is limited when the information cannot be linked to another dataset.

A basic example is the word “rainfall”. One system might record millimetres per day at a fixed weather station. Another might provide gridded satellite estimates, while a third stores a farmer’s manual observation as free text. If the date, location, unit and collection method are absent, analysts cannot confidently compare the records. A standard metadata record turns an isolated number into evidence with context.

The IASSIST perspective was especially relevant because research data specialists have long dealt with the problems of description, discovery, preservation and responsible reuse. Agriculture brings those problems into a high-volume, commercial and environmental setting. Data must work for a grower making a decision during a narrow weather window, for a researcher modelling food security over decades and for a policymaker monitoring land and water resources.

This also explains why standardisation is broader than choosing a file format. A comma-separated values file can be opened almost anywhere, yet its columns may still be ambiguous. A durable agricultural data system needs an agreed vocabulary, identifiers for places and organisations, documented methods, quality information and a way to express relationships between datasets.

Metadata turns measurements into usable evidence

Metadata is often treated as administrative material added after the “real” data has been collected. In digital agriculture, it is part of the evidence itself. A useful record might identify the crop variety, planting date, sensor model, sampling depth, coordinate reference system, weather conditions and processing steps. These details determine whether a result can be interpreted or reproduced.

Established approaches such as Dublin Core can support broad resource discovery, while DataCite-style metadata and persistent identifiers help identify datasets, reports and software. Geospatial standards such as ISO 19115 and services associated with the Open Geospatial Consortium are important when information is tied to paddocks, catchments, remote-sensing imagery or climatic zones. The precise standard matters less than consistent implementation and clear documentation.

Controlled vocabularies are equally important. “Wheat”, “Triticum aestivum” and a local shorthand may refer to related concepts without being interchangeable in every context. A machine-learning model trained on one classification scheme can produce misleading results when another system uses different crop categories. A shared vocabulary, with definitions and mappings between terms, allows researchers to compare data without erasing useful local detail.

Provenance adds another layer. Users need to know whether an observation came directly from a sensor, was edited by a technician, inferred from imagery or generated by a model. A yield map that has been smoothed, resampled and merged with soil data is different from the original harvester output. Recording those transformations supports accountability and helps future users judge whether the dataset is suitable for a new purpose.

Interoperability connects farms, laboratories and public agencies

Interoperability means that systems can exchange and interpret information without requiring every participant to use the same software. This is essential in a sector where producers may purchase machinery from several manufacturers, researchers work across institutions and public agencies collect data for different regulatory purposes. Open APIs, documented schemas and machine-readable formats can reduce the friction between those environments.

The distinction between syntactic and semantic interoperability is useful. Syntactic interoperability concerns the structure of a file or message: fields, encoding and data types. Semantic interoperability concerns what those fields mean. Two systems may both contain a field called “area” while one records hectares and the other records square metres. A technically successful data transfer can still be substantively wrong.

Deep learning, one of the subjects associated with the IASSIST 2017 programme, made this issue more urgent. Models can process enormous collections of images and sensor readings, but they inherit the weaknesses of their training data. If images are labelled inconsistently, if failed sensors are mistaken for dry conditions or if observations are concentrated in one region, the model may appear accurate while performing poorly elsewhere.

Australian conditions make this particularly clear. A system trained on dense, regular observations from a European cropping district may not translate neatly to broadacre farms in Western Australia or Queensland, where field sizes, soil types, rainfall patterns and connectivity differ sharply. A producer near Wagga Wagga may combine machinery data with local agronomy advice, while a station in the Northern Territory may rely on intermittent connectivity and satellite links. Interoperable systems must allow for these realities rather than assuming a single farm model.

Interoperability also has a commercial dimension. Farmers need confidence that data collected by one vendor can be exported, retained and used with another service. Where data is locked inside proprietary platforms, producers may lose control over their operational history. Standards cannot settle every question about ownership or competition, but portable formats and transparent interfaces can make switching and independent analysis more practical.

Governance determines who can use agricultural data

Technical standards do not decide whether data should be shared. Agricultural information can reveal business performance, land use, biosecurity risks, market position or culturally significant knowledge. A robust data framework therefore needs governance rules covering consent, access, licensing, security, retention and permitted reuse.

Public data and commercial data often have different purposes. A government agency may publish climate or land-use information under an open licence, while a farmer may treat yield maps and input records as confidential commercial assets. A university project may hold data under an agreement that permits research but restricts redistribution. Metadata should make these conditions visible instead of leaving users to infer them from a website or email exchange.

Australian data practice also needs to recognise Indigenous data sovereignty. Information connected to Country, traditional knowledge, cultural heritage or Indigenous communities cannot be governed adequately by technical access controls alone. The CARE principles—Collective Benefit, Authority to Control, Responsibility and Ethics—offer a useful complement to FAIR principles, which emphasise that data should be findable, accessible, interoperable and reusable. The two approaches address different responsibilities and should not be collapsed into a single checklist.

Local institutions illustrate the scale of the issue. CSIRO, state departments, universities, agribusinesses and organisations such as ABARES may all work with agricultural and environmental datasets, while water information can involve the Murray–Darling Basin and multiple jurisdictions. Their records may have different mandates, release schedules and levels of geographic detail. Shared governance profiles and persistent identifiers can clarify who created a dataset, who is accountable for it and which communities should be involved in decisions about reuse.

Good governance also supports trust at farm level. Australian producers are used to practical language—“the paddock”, “the block”, “the arvo weather window”—rather than abstract data-policy terminology. If a platform explains plainly what is collected, who sees it and how it benefits the business, adoption is more likely. Consent and licensing arrangements should be understandable to a grower, not written only for a legal or technical audience.

Global food security needs comparable data

The food-security discussions associated with IASSIST 2017 show why local agricultural records have value beyond an individual farm. Analysts study production, prices, nutrition, climate, trade, land use and supply-chain disruptions together. This requires data that can be compared across places and years without hiding important differences.

Comparable data does not mean identical data. A rainfall measure from a dryland wheat region should retain its local context, just as livestock data from northern Australia must reflect extensive grazing systems rather than intensive feedlots. Standards should preserve regional detail while making core elements consistent: location, time, unit, method, uncertainty and provenance.

Climate change raises the stakes. Heatwaves, droughts, floods and shifting seasons can affect production and food prices across connected regions. Researchers need historical observations, remote sensing, farm records and climate projections that can be aligned. If a dataset lacks a reliable coordinate reference, uses undocumented crop categories or changes its measurement method without notice, long-term analysis becomes fragile.

Australian examples are especially relevant because the market combines sophisticated export agriculture with large distances and uneven digital infrastructure. Grain, beef, cotton, horticulture and wine move through different supply chains, while regional businesses may face patchy broadband and limited technical support. A digital agriculture standard that works in a well-connected research station must still accommodate a producer managing data from a far-western property or an orchard operating with several seasonal contractors.

Standards also help connect farm data with consumer and environmental questions. A retailer may want traceability, a regulator may need chemical-use records, a researcher may model water demand and a community may care about biodiversity. These users do not need every raw record, yet they do need reliable summaries whose definitions and limitations are visible. Layered metadata and role-based access can support different uses without making all information public.

The lasting lesson from the conference archive

The archived IASSIST 2017 material presents data as a social and institutional resource, rather than a neutral pile of digital files. That framing is important for digital agriculture. Sensors and algorithms may produce the data, but people decide what to measure, how to describe it, who may access it and which risks are acceptable.

The conference themes also point towards a practical hierarchy. First, organisations need clear definitions for agricultural concepts and measurements. Next, they need metadata that records context, provenance and rights. They then need interoperable services and identifiers so that information can travel between platforms. Finally, they need governance that respects commercial confidentiality, public accountability and Indigenous authority.

This hierarchy is useful because many digital agriculture projects begin with the most visible technology: a dashboard, drone, image model or predictive service. Standards work is less visible, yet it determines whether the project can survive a change of vendor, staff member or research grant. A well-described dataset can be discovered years later; an undocumented export may become unusable as soon as its original software disappears.

For Australia, the practical test is whether these principles work across the whole sector: in a research laboratory, on a family farm, in a large corporate operation, through a government monitoring programme and along an export supply chain. The language may vary from “metadata” and “interoperability” in a conference session to “will this still work next season?” in the paddock. The underlying need is the same: trustworthy information that can move between people, systems and places without losing its meaning.

At the Conference

What attendees experienced in Lawrence

p1 reed

Plenary Sessions

Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.

platinum icpsr

Workshops

Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.

gold rockhurst

Social Events

An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.

Venue & Accommodations

Where the conference took place

silver ifdo

Kansas Union

University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.

silver ciser

The Eldridge

701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.

platinum icpsr

Springhill & TownePlace Suites

Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.

Program Highlights

Sessions and activities

Plenary Sessions Workshops Poster Presentations Committee Meetings Opening Reception Banquet Tour Kansas City Pecha Kucha Check-In Local Favorites

Getting Here

Lawrence, Kansas

Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045

Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).

Plan Your Stay

Accommodation options that were available

silver ifdo

The Eldridge

701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.

silver ciser

Springhill Suites

Marriott property. Room block reserved under "KU IASSIST Conference."

gold rockhurst

TownePlace Suites

Marriott property. Room block reserved under "KU IASSIST Conference."