Big data in the paddock: monitoring global crop health from space to soil
Across the wheatbelt of Western Australia and the heavy black soils of the Darling Downs, growers have long made sense of a season with a mix of gut feel, neighbourly chat, and the regular visit of the local agronomist. That recipe is being rewritten. Satellites pass overhead every few days, soil-moisture probes chatter away under the surface, and machine-learning models turn terabytes of raw signal into weekly dashboards that farmers check on their phones before smoko. The shift is part of a wider conversation playing out in forums such as IASSIST 2017, where researchers, data stewards, and policy makers are working through how best to archive, share, and reuse the torrent of agricultural information now flowing from the land.
Treating crop monitoring as a big-data problem is more than a technical exercise. A single hectare of wheat generates more measurements in a growing season than a farmer in the 1980s would have recorded across a working life. Multiply that by the roughly 22 million hectares sown to grain across Australia each year and then by hundreds of millions of hectares elsewhere, and the analytical challenge becomes obvious. Understanding crop health at that resolution feeds directly into food-security planning, commodity forecasting, and the quiet decisions that determine whether a family-run farm in the Riverina buys another header or sells off a block.
From observation to algorithm
The history of crop monitoring is a story of instruments getting sharper and datasets getting fatter. Visual field scouting still matters, but it samples only what a person can walk past. Aerial photography extended the view after the Second World War, then Landsat in the 1970s gave the world the first routine look at paddock-scale vegetation from orbit. Each step multiplied the amount of information a researcher had to handle.
The current wave is different in kind, not only in scale. Modern constellations such as Sentinel-2 and the commercial Planet fleet return imagery at three-to-five-metre resolution on a five-day revisit. Hyperspectral sensors break the light coming off a crop into hundreds of narrow bands, revealing stress from nitrogen deficiency or fungal infection days before the human eye notices a colour change. Drones, tractor-mounted cameras, and in-canopy sensors fill in the gaps at sub-metre scales. When these streams are fused with weather reanalyses, soil maps, and historical yield records, the resulting dataset runs into petabytes. The intellectual move has been to treat each pixel, each probe reading, each yield sample, as a row in a single, tidy table on which a model can be trained.
The data stack behind every paddock
A working crop-monitoring system rests on four layers, and Australian operators have come to refer to them as the data stack. The bottom layer is acquisition: satellites, drones, weather stations, and the simple act of a grower uploading a header yield map. The second is curation, where raw files are georeferenced, quality-checked, and stored in archives that researchers can actually find five years later. The third is modelling, where statistical and machine-learning methods turn inputs into forecasts of biomass, yield, or disease risk. The top layer is delivery: dashboards, SMS alerts, or paper reports tailored to a farm manager, a trading desk, or a federal policy unit.
Where the layers connect matters as much as what sits in each one. A radiometric index such as NDVI is only useful if the imagery behind it has been atmospherically corrected against a trustworthy surface-reflectance product. A yield forecast only earns credibility if the weather inputs come from a reanalysis that captures the local quirks of, say, a sea breeze rolling inland from Spencer Gulf. The quiet craft of data stewardship, the same craft that IASSIST conferences have championed for decades, is what turns a pretty map into a decision a grower can defend to a bank manager.
Lessons from Australia's grain belt
Australia's grain belt offers a useful natural laboratory because the climate refuses to behave. El Niño and La Niña cycles swing rainfall by hundreds of millimetres between seasons, and the line between a record harvest and a failed one can come down to a single October rain band. Over the past decade, public and private partners have layered big-data tools onto this volatility with notable results. CSIRO and the Grains Research and Development Corporation have invested in analytics platforms that combine satellite imagery, APSIM crop modelling, and on-farm yield data. CBH Group, the grower-owned cooperative that handles the bulk of Western Australia's harvest, runs its own seasonal forecasting pipeline to plan receivals and shipping out of Kwinana, Esperance, and Geraldton.
The table below sketches how several monitoring approaches stack up against one another when judged on the realities of Australian broadacre farming.
| Method | Spatial resolution | Typical lag | Capital cost | Best suited to |
|---|---|---|---|---|
| Sentinel-2 multispectral | 10–20 m | 3–5 days | Low (free imagery) | Regional biomass and nitrogen status |
| Commercial very-high-resolution | 0.3–3 m | Daily tasking possible | High per scene | Paddock-scale disease scouting |
| Drone multispectral | 1–5 cm | On demand | Moderate hardware, low per flight | Trial plots and patchy stress |
| Tractor-mounted optical sensors | Sub-metre, in-row | Real time | Bundled with new machinery | Variable-rate nitrogen application |
| Soil-moisture probe networks | Point, representative | Hourly | Moderate, ongoing telemetry | Irrigation scheduling in the Murray–Darling |
Two patterns stand out from Australian experience. First, no single source covers everything; a Sunraysia table-grape operation and a Wimmera wheat grower need different mixes of the stack. Second, the value of a dataset rises sharply once it has been collected consistently for several seasons, because the long tail of historical records is what lets a model detect that this year is genuinely different rather than simply noisy.
Reading the signals from NDVI to nitrogen
Turning pixels into advice is where analytics earns its keep. NDVI, the normalised difference vegetation index, is the workhorse because it is easy to compute and tolerates a wide range of sensors. More advanced indices, such as the chlorophyll absorption ratio index or thermal-derived canopy temperature, pick up on the subtle changes that precede visible stress. When these are fed into machine-learning models trained on past seasons, they can forecast yield weeks before harvest and flag individual paddocks that are likely to fall short of the budget used to plan sowing, fertiliser, and forward sales.
The next step is linking those forecasts to management. Variable-rate application maps, written from satellite and sensor data, allow a grower to apply nitrogen only where it will pay back, sparing both the budget and the local catchment. In the Murrumbidgee and Murray valleys, water authorities and irrigators have begun using similar analytics to allocate limited water against the crop responses it is most likely to produce. The aim is not to replace agronomic judgement but to give it a sharper instrument. A seasoned agronomist still walks the crop, but now walks it with a tablet showing where the model expects trouble and where it does not.
Sharing the harvest through open data and governance
Big-data crop monitoring only delivers public-good outcomes if the underlying information can flow. That means dealing with questions of licensing, privacy, and provenance that are familiar to anyone working with administrative or health data. A paddock yield map can reveal a farm's commercial position; a soil-carbon measurement can affect the value of a carbon credit. Stewardship frameworks developed for census and clinical data are being adapted to agriculture, with trusted data repositories, persistent identifiers, and clear rules about who can access what and on what terms.
International collaboration is already reshaping the picture. Groups coordinated through the Group on Earth Observations, the CGIAR platform for big data in agriculture, and the Food and Agriculture Organization of the United Nations are knitting together national archives so that a researcher in Narrabri can compare notes with one in Nairobi or Saskatchewan. Standards such as STAC for spatiotemporal catalogues and DataCite for citation are making it easier to reuse datasets years after collection. For Australia, the strategic prize is twofold: better national preparedness for climate-driven shocks, and a stronger voice in the global analytics that increasingly set the price of the grain leaving Portland or Newcastle.
Practical guidance for growers, researchers, and data stewards
- Start with the agronomic question, not the sensor. Choose a method that matches a decision you would otherwise make on instinct, and only then pick the data stream that informs it.
- Insist on documented provenance. Every dataset used in a model should carry metadata about who collected it, when, with what instrument, and under what licence.
- Plan for the long tail. A monitoring system that runs for two seasons is a pilot; one that runs for ten is a research platform. Storage, governance, and funding should be sized for the long game.
- Combine scales deliberately. Pair satellite imagery with point sensors and yield data so that broad patterns are anchored in ground truth.
- Build trust through transparency. Document model assumptions and limitations openly, and publish validation results alongside forecasts.
- Engage with national archives. Deposit well-documented datasets with bodies such as CSIRO Data Access Portal or university repositories so future researchers can find them.
- Treat data stewardship as core infrastructure. Allocate staff time and budget for curation, not just acquisition and modelling.
At the Conference
What attendees experienced in Lawrence
Plenary Sessions
Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.
Workshops
Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.
Social Events
An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.
Venue & Accommodations
Where the conference took place
Kansas Union
University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.
The Oread
1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.
The Eldridge
701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.
Springhill & TownePlace Suites
Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.
Program Highlights
Sessions and activities
Getting Here
Lawrence, Kansas
Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045
Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).
Plan Your Stay
Accommodation options that were available
The Eldridge
701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.
The Oread
1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.
Springhill Suites
Marriott property. Room block reserved under "KU IASSIST Conference."
TownePlace Suites
Marriott property. Room block reserved under "KU IASSIST Conference."