Abstract low-poly geometric background in shades of pale blue and white

Data in the Middle: The common language of research

Plenary Ideas From IASSIST 2017

IASSIST 2017, held at the University of Kansas in Lawrence from 23–26 May 2017, examined how research data can function as a shared language across disciplines. Its theme, “Data in the Middle: The Common Language of Research,” placed data professionals between researchers, institutions, policymakers and the public. The archived programme brings together plenary talks, specialist sessions and practical discussions about managing information throughout the research lifecycle.

The plenary presentations are especially useful for understanding the conference’s wider argument. Big data, deep learning, digital agriculture and global food security were treated as connected areas rather than isolated technology trends. For Australian researchers and information professionals, the material remains relevant because it addresses familiar questions around drought monitoring, public-sector datasets, reproducibility, data stewardship and responsible reuse.

Presentation focus Main issue Continuing relevance
Big data and research infrastructure How can large, complex datasets be made useful and trustworthy? Supports national research infrastructure and cross-institutional collaboration
Deep learning and machine intelligence What can automated methods reveal, and what must researchers still explain? Important for health, climate, agriculture and social research
Digital agriculture How should field, satellite and sensor data be curated? Applies to Australian farming, water management and land-use studies
Global food security How can data support decisions across regions and communities? Connects local evidence with international development and policy
Data stewardship Who preserves, describes, governs and enables access to research data? Relevant to FAIR data, Indigenous data governance and public accountability

Data As A Shared Research Language

The conference theme proposed that data has a mediating role. It sits between an observation and an argument, between a research team and a repository, and between a specialist method and a public decision. A dataset is therefore more than a collection of values. Its meaning depends on documentation, definitions, provenance, permissions and the circumstances in which it was created.

This perspective helps explain why the plenary programme moved across disciplines. A social scientist, agricultural researcher, statistician and data librarian may use different terminology while working with similar underlying concerns: how variables are defined, how records are linked, how uncertainty is communicated and how future users can interpret the material. The presentations encouraged attendees to see data curation as an intellectual activity rather than a final administrative step.

For Australian institutions, that message fits the structure of the local research environment. Universities in Sydney, Melbourne, Brisbane and Perth collaborate with government agencies, hospitals, industry and community organisations. The Australian Bureau of Statistics, state departments and national research facilities all publish or manage data with different standards and access conditions. A common language can help these groups negotiate consistent metadata, ethical safeguards and preservation responsibilities.

The archive also reflects the practical culture of research support. Check-in details, transport information and local recommendations appear alongside scholarly sessions, reinforcing the idea that conferences create working relationships as well as distribute information. The human network around a dataset often determines whether a valuable collection is reused or quietly becomes inaccessible.

Big Data Requires More Than Scale

The big-data discussions focused on the gap between volume and value. Large datasets can reveal patterns that smaller studies miss, yet size creates problems involving storage, computation, quality control and interpretation. A million records do not automatically provide a more reliable answer than a carefully designed sample. The central task is to understand what the data represents and where its limits lie.

The plenary perspective placed infrastructure and method together. Researchers need systems that can ingest, preserve and process information, but they also need clear records of collection methods and transformations. A model trained on inconsistent categories can produce precise-looking results that remain conceptually weak. For data services, this means retaining codebooks, version histories, processing scripts and links between raw and derived files.

That lesson is important in Australia’s policy environment. Public agencies increasingly release datasets through national and state portals, while the Privacy Act 1988 places obligations around personal information. De-identification is not a magic solution: location, dates and combinations of seemingly harmless variables can make individuals or small communities recognisable. Good stewardship therefore combines technical controls with risk assessment, consent practices and transparent access procedures.

Everyday data habits also shape the issue. Australians routinely use transport apps, loyalty programmes, online banking and health portals, creating detailed digital traces. Researchers may gain valuable evidence from these sources, but responsible reuse requires more than obtaining a file. Clear purpose, lawful handling, proportionality and explanation are necessary if public trust is to survive beyond a single project.

Deep Learning And The Role Of Interpretation

The deep-learning strand presented machine intelligence as a powerful research method rather than an independent substitute for scholarship. Neural networks can classify images, detect patterns in text, estimate outcomes and process signals at a scale that would overwhelm manual analysis. Their usefulness grows when they are paired with carefully prepared training data and a research question grounded in domain knowledge.

The presentations also invite scrutiny of the assumptions behind automated analysis. A model can inherit historical bias, perform poorly when conditions change or identify correlations without explaining causation. Researchers must record the source population, labelling process, model version, evaluation measures and known limitations. Reproducibility depends on these details, even when the complete dataset cannot be shared.

This concern has practical force in Australian settings. An image-recognition system trained on North American crops may struggle with Australian varieties, harsh sunlight or drought-affected fields. A language model built from metropolitan news may misread regional communities or Aboriginal and Torres Strait Islander contexts. Local validation, representative data and consultation with affected groups are essential before automated tools are used in public decisions.

The plenary message was therefore balanced. Deep learning expands what researchers can examine, but it does not remove the need for librarians, archivists, statisticians, subject experts or communities. Their work supplies the context that makes an output meaningful. Documentation should explain how a result was generated, while governance should establish who may challenge, correct or withdraw it.

Agriculture, Food Security And Data Curation

Digital agriculture provided one of the clearest bridges between technical innovation and social consequence. Farms now generate information through weather stations, satellite imagery, machinery, soil sensors, yield monitors and farm-management software. When joined with historical climate records and market information, these sources can support decisions about planting, irrigation, disease control and supply chains.

The conference’s agricultural discussions treated curation as a condition of usefulness. Different projects may describe the same field, crop or event in incompatible ways. Sensor readings can contain gaps, changing calibration and uncertain location information. Researchers must preserve the original observations while documenting cleaning, interpolation and aggregation. Without that record, later users cannot distinguish a real environmental pattern from an artefact of processing.

The archived discussion of agricultural data curation extends this point by focusing on the difficulties of working with large-scale studies. Its relevance is immediate in Australia, where drought, bushfire, salinity and water allocation produce long-running research needs. A dataset collected during a dry season may answer a different question from one gathered after heavy rain, even when the fields and instruments appear identical.

Food-security research also requires attention to geography and inequality. A national average can hide the experience of remote communities, small producers or households affected by rising prices. Data from supermarkets and wholesale markets may show supply movements, while community-level evidence explains access and affordability. In Australian cities, household food choices are shaped by supermarket concentration, transport costs and the distance between outer suburbs and fresh-food outlets.

These issues make interoperability a social question as well as a technical one. Researchers need shared vocabularies, stable identifiers and open standards, yet communities should retain a say in how knowledge about land, culture and resources is collected and reused. Indigenous data sovereignty principles are particularly important where datasets concern Country, cultural knowledge or community wellbeing.

Preservation, Access And Responsible Reuse

The final theme running through the plenary programme is the long-term life of research data. A presentation can end when the conference session closes, but a dataset may need to remain interpretable for decades. Preservation involves file formats and storage, while stewardship also includes rights information, discoverability, documentation and plans for future migration.

The archive’s programme is valuable precisely because it demonstrates this longer life. Conference materials that were created for a particular week in 2017 can still support professional learning when they retain their context. Schedules, presentation descriptions and practical information help users understand what was discussed, who the intended audience was and how the sessions fitted together. Archiving is therefore an act of interpretation as well as retention.

For Australian practitioners, this connects with open research expectations, grant requirements and institutional repositories. A project may need to provide a data-management plan, deposit suitable outputs and explain restrictions on sensitive material. The Australian market also contains commercial data providers whose licensing terms can prevent redistribution, even when researchers have paid for access. Reuse depends on reading contracts carefully and recording what future users are permitted to do.

International professional networks continue to carry this work forward. The IASA conference offers a useful point of comparison for people interested in statistical computing, data methods and cross-disciplinary exchange. Events of this kind show that technical knowledge develops through communities that share standards, examples and critical practice.

For readers revisiting IASSIST 2017, the most durable lesson is that data management cannot be separated from research design. Decisions about collection, description, access and preservation affect the claims that can later be made. The plenary presentations encourage a form of stewardship that is technically capable, ethically alert and attentive to the people represented in the data.

Practical Lessons For Australian Data Professionals

The plenary ideas can be translated into everyday decisions by librarians, research-office staff, analysts, archivists and investigators. The following priorities reflect the conference’s emphasis on data as a shared, durable resource:

  • Document meaning before scale: Record definitions, units, collection conditions, sampling decisions and known gaps before combining large datasets.
  • Build provenance into workflows: Preserve original files, code, model versions and transformation histories so that results can be checked or reproduced.
  • Assess privacy and community impact: Apply the Privacy Act 1988 where relevant, and use culturally appropriate governance for Indigenous and community-held information.
  • Design for Australian conditions: Test models and metadata against regional, rural, remote and metropolitan contexts rather than assuming overseas training data will transfer cleanly.
  • Separate access from openness: Make suitable material discoverable while using controlled access, licensing and secure environments for sensitive or commercially restricted data.
  • Plan for future interpretation: Deposit clear documentation with the data, identify responsible custodians and choose formats that can be maintained after the original project ends.

IASSIST 2017 remains a useful reference because its plenary presentations connected emerging technologies with the quieter work that makes research dependable. Big data and deep learning attracted attention, yet their value depended on curation, governance and collaboration. The same principle applies to digital agriculture and food-security research: reliable evidence begins with careful stewardship and ends with responsible interpretation.

At the Conference

What attendees experienced in Lawrence

p1 reed

Plenary Sessions

Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.

platinum icpsr

Workshops

Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.

gold rockhurst

Social Events

An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.

Venue & Accommodations

Where the conference took place

silver ifdo

Kansas Union

University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.

silver ciser

The Eldridge

701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.

platinum icpsr

Springhill & TownePlace Suites

Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.

Program Highlights

Sessions and activities

Plenary Sessions Workshops Poster Presentations Committee Meetings Opening Reception Banquet Tour Kansas City Pecha Kucha Check-In Local Favorites

Getting Here

Lawrence, Kansas

Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045

Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).

Plan Your Stay

Accommodation options that were available

silver ifdo

The Eldridge

701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.

silver ciser

Springhill Suites

Marriott property. Room block reserved under "KU IASSIST Conference."

gold rockhurst

TownePlace Suites

Marriott property. Room block reserved under "KU IASSIST Conference."