Abstract low-poly geometric background in shades of pale blue and white

Data in the Middle: The common language of research

Bridging Silos: Integrating Food Security Data Across Research Partners

Food security research has shifted decisively from solo labs to sprawling consortia. A single project may now bring together soil scientists, economists, hydrologists, public health researchers and remote sensing analysts, each working from their own institutional vantage point. The promise is real: richer models, faster learning cycles, and policy advice that reflects the full complexity of how people eat, farm and trade. The friction is just as real. When every partner arrives with their own datasets, schemas, vocabularies and storage habits, the act of stitching everything together can quietly become the most expensive line item in the whole program.

That friction is especially visible in multi-institutional food security work, where the same word might mean different things to an agronomist in Narrabri and a trade modeller in Canberra. Drawing on themes raised at the IASSIST 2017 gathering on shared research data, this piece looks at the integration challenges that crop up again and again, and at the practical choices that help projects move past them.

Why food security projects multiply the integration problem

Food security is not a single question; it is a stack of questions. Yield forecasting, household consumption, market prices, climate shocks, water availability, nutrition outcomes and policy responses all sit inside the same conversation. Each of those strands tends to live with a different research group, often in a different university or agency, and almost always with a different funding cycle attached. The data integration challenge is therefore not just technical; it is the visible edge of a much larger organisational puzzle.

When a consortium spans continents, the variance grows. A team in Adelaide tracking wheat trials, a counterpart in Nairobi monitoring maize markets and a satellite data unit in Toulouse all bring assets that look incompatible on the surface. Field books are paper notes in one location, standardised spreadsheets in another, and proprietary sensor outputs in a third. Layer in language, units (tonnes versus bushels, hectares versus acres) and timing conventions, and the integration work multiplies well beyond a simple file merge.

The Australian context sharpens this further. Researchers working across the Murray–Darling Basin are used to juggling state-level land-use records, Commonwealth water accounting and farm-level data that owners will only release under strict conditions. Add an international partner looking at South-East Asian rice systems, and the translation work doubles overnight. The cost of doing nothing is high: silos leave insights trapped inside institutions, duplicating effort and quietly eroding public trust when numbers published by different agencies disagree.

Common friction points across institutions

The first friction point is almost always vocabulary. One partner's "smallholder" might exclude peri-urban market gardeners; another partner's might include them. Without a shared semantic layer, queries return confusing results and analysts spend their early weeks debating definitions instead of testing hypotheses. A related problem is time-zone and date handling, particularly when field observations are recorded locally but reported against international reporting periods.

A second friction point is data provenance. Multi-institutional food security projects often inherit legacy datasets whose lineage is partially known at best. Was that rainfall figure recorded by a gauge or estimated from satellite? Was household income measured before or after a particular subsidy? Without provenance captured in a structured way, integrated datasets can quietly amplify bias, and reproducibility becomes a polite fiction.

A third friction point is access control. A surprising number of projects stall not because the data cannot be moved, but because legal and ethical frameworks forbid it. Sensitive farm business data in Australia is one example: growers may consent to share with a university, yet balk at onward sharing with international partners or commercial analysts. These constraints are reasonable, and any integration plan that ignores them tends to fail in front of a steering committee.

Finally, infrastructure rarely lines up. Some groups work on modestly resourced servers, others push terabytes through high-speed research networks. Tools that run beautifully in a well-funded CSIRO cluster can crawl on a regional university system. The integration layer has to be designed for the slowest link, not the fastest.

Standards and metadata as the first stabiliser

Metadata is not glamorous, but it is the cheapest insurance a project can buy. When every contributor describes their datasets with the same schema, with controlled vocabularies, units and temporal coverage, downstream integration stops being archaeology and starts being engineering. Standards such as the FAIR principles — findable, accessible, interoperable and reusable — provide a checklist even when full interoperability is unrealistic in year one.

Controlled vocabularies are particularly valuable. AGROVOC, the Food and Agriculture Organization's thesaurus, offers one shared reference for crop, livestock and food system terms. Pair it with a domain-specific extension for regional realities, including indigenous food names, local market categories and Australian crop codes, and you get a workable translation layer. Schema.org's Dataset markup helps datasets surface in search tools, which sounds trivial but matters enormously when new partners join mid-project and need to find what already exists. Dublin Core still has a place as a lightweight, low-friction metadata profile for small contributors who cannot afford heavy curation. The right choice depends on the project's lifespan, partner mix and ambition, as the comparison below suggests.

Standard Strengths Limits Best fit for
Dublin Core Simple, low overhead, easy for non-specialists to adopt Shallow semantics, weak for quantitative datasets Light metadata needs in survey projects
AGROVOC Domain-specific, multilingual, FAO-backed Requires mapping to local terms, governance overhead Cross-border food systems research
DataCite + schema.org/Dataset Excellent discoverability through search and DOI registration Light on domain semantics Open publication of finished datasets
ISO 19115 (geographic) Rich spatial and temporal metadata Heavy, requires trained cataloguers Remote sensing and land-use layers
CF/ACDD (climate/netcdf) Native to climate and earth observation communities Steep learning curve outside those communities Observational networks and climate data

Even a partial move toward shared metadata delivers outsized benefits. A simple agreement that every dataset must carry contributor, location, units, time coverage and access conditions, even if the rest is free-form, removes a surprising amount of downstream pain.

Governance, trust and the human layer

The hardest part of data integration is rarely the data. It is the people. Researchers protect datasets they spent years collecting. Funders want credit lines in every paper. Ethics boards in different jurisdictions interpret consent differently. A project that treats governance as an afterthought will discover these tensions only when something breaks, usually at the worst possible moment.

A workable governance model usually has three layers. A high-level data sharing agreement covers the legal and ethical floor: who can use what, for what purpose, with what obligations. A data management plan, ideally written at the proposal stage and updated annually, captures operational decisions around storage, retention and documentation responsibilities. A working group with a rotating chair and clear escalation paths handles the day-to-day frictions. Without that middle layer, the high-level agreement becomes a paper artefact that nobody reads until there is a dispute.

Trust is built through small, repeated acts of competence. Acknowledging a partner's data in a paper is easy and often skipped. Cleaning a dataset you did not collect is tedious and quietly powerful. Responding quickly when a colleague flags a coding error builds faster than any memorandum of understanding. Australian projects that bring together state departments, the CSIRO and grower bodies have learned that trust accrues from concrete outputs, not from signed agreements alone. There is a useful local analogy. Australians grow up learning that a backyard barbie is less about the food and more about the conversation, so it is with multi-institutional data work: integration succeeds when contributors share the same fire and a willingness to keep turning the snags.

Working examples from Australian collaborations

A handful of recent Australian collaborations illustrate how the pieces fit together. The Grains Research and Development Corporation's investments in digital agriculture have pushed partners toward common yield reporting formats, partly because the alternative was publishing contradictory numbers during a politically charged drought conversation. The lessons learnt there flow back into national policy advice on water availability in the Murray–Darling Basin.

University-led initiatives have shown what shared infrastructure can do. A consortium involving the University of Queensland, the University of Sydney, RMIT and several state agencies built a federated data layer for tropical food security research spanning northern Australia and South-East Asia. Rather than pooling everything in one warehouse, each partner kept local control while agreeing on metadata, access windows and analysis environments. The federation model kept the trust intact and the data local.

The lessons are not all technical. Projects that invested in a dedicated data steward — a real person, not a job description — reported fewer integration disputes and faster onboarding for new partners. Projects that relied on a rotating PhD student to "do the data stuff" reported the opposite. Funding bodies in Australia are starting to notice, and several now ask explicitly for data stewardship in funding applications.

Bushfire response work has also accelerated practical thinking. After the 2019–20 season, researchers coordinating recovery studies found themselves trying to merge datasets collected by farmer groups, state agencies, NGOs and Commonwealth bodies in the middle of an emergency. The experience pushed several groups toward pre-positioned data agreements and lightweight shared schemas that could be activated quickly. That kind of preparation, unglamorous and easy to defer, is what separates an integrated response from a chaotic one.

Practical pathways toward durable integration

Most projects can move from chaos to workable integration in three realistic stages. The first is alignment, agreeing on a small, finite set of shared metadata fields, vocabularies and access tiers before any data is moved. The second is federation, building an integration layer that respects local control while enabling joint analysis. The third is curation, an ongoing investment in data stewards, documentation and quality assurance long after the launch press release has faded.

Recommendations worth borrowing from projects that have already done the hard yards:

  • Pick a metadata schema shared by every collaborator, even if imperfect, before designing the integration architecture.
  • Treat provenance as a first-class output: capture who, when, how and under what consent for every record.
  • Build a federation layer rather than a central warehouse where legal, ethical or commercial constraints are present.
  • Fund a dedicated data steward for the life of the project, not just the build phase.
  • Run quarterly integration health checks using a small set of agreed indicators, such as schema compliance rates and median time to resolve a flagged issue.
  • Publish a plain-language data catalogue for new collaborators, decision-makers and the curious public.
  • Schedule a "data sunset" review at project end, so findings remain reusable long after the funding cycle closes.

The temptation in food security work is to treat data integration as plumbing: necessary, unglamorous, best left to the technicians. That view is expensive. Integration choices shape what questions can be asked, which partners can join, and whether the work survives the next drought, the next funding round or the next leadership change. Treat the work as craft, not cabling, and the consortia that emerge will be the ones worth keeping.

At the Conference

What attendees experienced in Lawrence

p1 reed

Plenary Sessions

Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.

platinum icpsr

Workshops

Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.

gold rockhurst

Social Events

An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.

Venue & Accommodations

Where the conference took place

silver ifdo

Kansas Union

University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.

silver ciser

The Eldridge

701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.

platinum icpsr

Springhill & TownePlace Suites

Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.

Program Highlights

Sessions and activities

Plenary Sessions Workshops Poster Presentations Committee Meetings Opening Reception Banquet Tour Kansas City Pecha Kucha Check-In Local Favorites

Getting Here

Lawrence, Kansas

Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045

Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).

Plan Your Stay

Accommodation options that were available

silver ifdo

The Eldridge

701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.

silver ciser

Springhill Suites

Marriott property. Room block reserved under "KU IASSIST Conference."

gold rockhurst

TownePlace Suites

Marriott property. Room block reserved under "KU IASSIST Conference."