Abstract low-poly geometric background in shades of pale blue and white

Data in the Middle: The common language of research

Bridging the gap: lessons from IASSIST 2017 on data sharing

The archive of the IASSIST 2017 conference, hosted by the University of Kansas in Lawrence from 23 to 26 May 2017, remains a valuable resource for those studying how research communities translate raw information into knowledge. The central theme, "Data in the Middle: The Common Language of Research," addressed the friction points that emerge whenever the people generating datasets must hand them over to people wanting to analyse, repurpose, or simply understand them.

Australia's research landscape stretches from the coral-reef monitoring projects of the Great Barrier Reef Marine Park Authority to the long-running agricultural trials at CSIRO sites in Canberra and Adelaide. Many of the datasets generated in these programs eventually pass through services funded by the Australian Research Data Commons (ARDC) or the now-concluded Australian National Data Service (ANDS). Practitioners who attended the Lawrence sessions often returned home to Brisbane and Perth with fresh ideas on how documentation, governance, and skill-building could be improved in their own workflows.

The conference offered not just theoretical reflections but practical examples of how libraries, archives, and domain specialists can work alongside scientists and policy officers. Several sessions highlighted machine learning pipelines, digital agriculture projects, and global food security initiatives, each presenting a different angle on the relationship between those who produce data and those who consume it. The remainder of this article explores what Australian practitioners can still take away from those exchanges.

Metadata as the shared vocabulary

One of the recurring themes during the plenary panels was that metadata, far from being a bureaucratic afterthought, acts as the connective tissue between data producers and data users. Without consistent descriptive frameworks, a dataset on wheat yield from a New South Wales field station is functionally inaccessible to a modeller working in Melbourne. Speakers at Lawrence stressed the importance of adopting community-recognised schemas such as Dublin Core, DataCite, and domain-specific standards used widely in the agricultural sciences.

The FAIR principles — findable, accessible, interoperable, and reusable — received substantial attention. Presenters argued that data producers often assume their preferred laboratory notebook conventions will travel with the dataset, while users arrive expecting standardised vocabularies, persistent identifiers, and machine-readable licences. Bridging the two requires deliberate curation rather than passive assumption, and several contributors warned that quality assurance at the documentation stage is often cheaper than retroactive cleanup.

Australian audiences particularly noted the parallels with the work of the National eResearch Collaboration Tools and Resources project (Nectar) and the long-running efforts of the Australian Bureau of Statistics to harmonise its statistical geographies. Both organisations have learned the hard way that metadata quality determines whether data lives or dies inside an institutional repository.

Trust, stewardship, and the human element

Beyond schemas and standards, conference attendees returned repeatedly to questions of trust. Data producers worry about misinterpretation, misuse of sensitive records, and the erosion of credit when their carefully curated files appear in derivative works. Data users, in turn, often suspect that the files they receive have been altered, simplified, or stripped of vital context. Both groups agreed that institutional reputation is built slowly and lost quickly when something goes wrong with a shared dataset.

Speakers from university libraries in North America and Europe presented stewardship models that explicitly address these anxieties. Curation as a service, embedded data champions within faculties, and rolling credit mechanisms emerged as common ingredients. The University of Kansas organisers pointed to mid-career data stewards as a profession in formation, one that combines librarianship, archival science, and applied statistics.

For Australian institutions such as the Australian National University or the University of Sydney, these models echo ongoing conversations about professional recognition for research data managers. The ARDC has invested heavily in capability uplift, and the lessons from IASSIST suggest that recognition of stewardship roles may matter as much as technical training when trust between producers and users needs to be reinforced.

Open data and its practical limits

The Lawrence programme included spirited discussion of open data as both an aspiration and an operational reality. Proponents cited benefits ranging from reproducible science to citizen engagement, while sceptics pointed to privacy risks, commercial sensitivities, and the genuine cost of preparing data for public release. Australian audiences recognised these tensions immediately, given the cultural sensitivities embedded in Indigenous Data Sovereignty work led by bodies such as the Maiam nayri Wingara Indigenous Data Alliance.

Several IASSIST presenters emphasised tiered access models, where publicly funded research increasingly lands in open repositories like the ARDC Data Portal while sensitive cohorts remain in controlled environments. The idea is not to choose between openness and restriction but to design release mechanisms appropriate to each dataset's provenance and risk profile.

Practical examples from digital agriculture sessions were particularly resonant for Australian readers. Open datasets on rainfall, soil moisture, and crop yields supported local farmer decision-making in the Murray-Darling Basin and international food-security modelling. The constraint, presenters agreed, lay not in principle but in the labour required to document, clean, and publish data at a quality users can actually trust.

Skills, training, and data literacy

A sizeable portion of the IASSIST 2017 schedule addressed training, with workshops on everything from introductory R and Python for new data users to advanced metadata curation for librarians. The implicit message was that bridging producer-user gaps requires sustained investment in human capacity, with technical infrastructure serving as the necessary but partial foundation.

Australian universities have long recognised this challenge. Initiatives such as the Adelaide-based Centre for Data and Knowledge Integration and various Graduate Research Schools in Melbourne and Brisbane have developed short courses aimed at postgraduate cohorts. The discussions in Lawrence encouraged these programmes to extend their reach beyond early-career researchers and include senior academics, policy advisers, and even media professionals who routinely interpret data for the public.

Participants also highlighted the value of cross-disciplinary micro-credentials, where social scientists, ecologists, and statisticians can share a common introduction to data handling. Such shared foundations mirror the conference's own format, where librarians, computer scientists, and domain specialists sat alongside one another to discuss what "good enough" data actually means in practice.

Ethics, provenance, and cultural context

Ethics surfaced repeatedly during the data-sharing sessions, particularly in panels touching on global food security. Presenters stressed that a dataset stripped of its geographic, temporal, and cultural context risks being misread by readers far from its origin. Several speakers cautioned against treating data as universally portable while ignoring how local knowledge conditions its interpretation.

This concern resonated strongly with Australian researchers working on Aboriginal and Torres Strait Islander data, where governance protocols require communities to retain control over how their information is used. The Maiam nayri Wingara principles of Indigenous Data Sovereignty align with the broader IASSIST conversation that ethics cannot be reduced to a checkbox exercise; it must inform collection, curation, and release decisions from the outset.

Conference recordings also touched on issues of consent in longitudinal studies, anonymisation in health datasets, and the politics of attribution when datasets circulate across borders. Each scenario reinforced the value of provenance metadata — the documentation that records who collected the data, under what circumstances, and with what expectations of future use.

Infrastructure, sustainability, and funding

Underpinning every other theme was the sober recognition that bridging data producers and users requires stable infrastructure. The Australian experience is instructive: the winding down of ANDS in 2018 and the consolidation of capabilities under the ARDC demonstrated both the strengths and the fragility of national data services. Funding cycles, presenter after presenter noted, often outlast the political attention given to research infrastructure.

Lawrence participants discussed cloud-based repositories, persistent identifier services, and the economics of long-term preservation. Several North American examples demonstrated how consortia pooling storage and curation effort could lower per-institution costs. Australian counterparts pointed to NCRIS-funded platforms like the Australian Imaging and Biomarker Resource and the Population Health Research Network as parallel examples of shared infrastructure lowering barriers for smaller universities and regional research hubs.

Sustainability also meant investing in people alongside platforms. The conference highlighted how data stewards, repository managers, and software engineers require career pathways that can survive changes in government and grant agency priorities. Without those pathways, even the best-designed infrastructure cannot reliably serve the data producers and users who rely on it.

Comparing the roles of data producers, users, and stewards

The table below summarises the primary concerns, common tools, and signature contributions of each stakeholder group, drawing on the cases discussed across the Lawrence programme and Australian examples such as the ARDC and CSIRO data initiatives.

Aspect Data producers Data users Data stewards
Primary concern Recognition, accuracy, downstream interpretation Access, documentation, comparability Preservation, interoperability, governance
Typical tools Laboratory information systems, field notebooks, sensors Analytical software (R, Python, SPSS), notebooks, dashboards Repositories (ARDC, ANDS, Dataverse), metadata editors, persistent identifier services
Skill emphasis Domain methodology, instrument calibration, preliminary documentation Statistical literacy, reproducibility, modelling Curation, metadata standards, legal and ethical frameworks
Australian example CSIRO digital agriculture teams in Adelaide and Canberra Researchers in Melbourne and Brisbane using the ARDC portal ARDC staff and university data managers carrying the ANDS legacy
Reward signal Publications, dataset citations, methodological contribution Insights, policy briefs, peer-reviewed findings Institutional recognition, professional certification, community reputation

The conference archive, still searchable through the University of Kansas hosts, leaves a clear lesson: sustainable data exchange depends on platforms and policies working in harmony with people. It depends on people who understand each other's constraints, who document their work generously, and who are rewarded for stewardship as much as for discovery. For Australian practitioners from Hobart to Perth, the IASSIST 2017 conversations remain a useful reference point whenever the practical work of connecting datasets with the communities that need them takes centre stage.

At the Conference

What attendees experienced in Lawrence

p1 reed

Plenary Sessions

Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.

platinum icpsr

Workshops

Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.

gold rockhurst

Social Events

An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.

Venue & Accommodations

Where the conference took place

silver ifdo

Kansas Union

University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.

silver ciser

The Eldridge

701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.

platinum icpsr

Springhill & TownePlace Suites

Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.

Program Highlights

Sessions and activities

Plenary Sessions Workshops Poster Presentations Committee Meetings Opening Reception Banquet Tour Kansas City Pecha Kucha Check-In Local Favorites

Getting Here

Lawrence, Kansas

Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045

Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).

Plan Your Stay

Accommodation options that were available

silver ifdo

The Eldridge

701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.

silver ciser

Springhill Suites

Marriott property. Room block reserved under "KU IASSIST Conference."

gold rockhurst

TownePlace Suites

Marriott property. Room block reserved under "KU IASSIST Conference."