Abstract low-poly geometric background in shades of pale blue and white

Data in the Middle: The common language of research

Best practices for storing and preserving conference datasets

Conference datasets are rarely static artefacts. They grow during the event itself, swell with supplementary slides, audio recordings and participant notes, and often take on a second life when they are cited, reused or reanalysed years later. The presentations shared at IASSIST 2017 in Lawrence demonstrated how quickly a single panel can yield dozens of files in mixed formats, and the organisers have since reflected on the practical lessons that emerged from that week. Researchers and event coordinators who treat these materials as a planned scholarly record, rather than as loose attachments to a programme, lay the groundwork for genuine long-term value.

Australian institutions sit at an interesting crossroads in this conversation. The Australian Research Data Commons has spent more than a decade building shared infrastructure that researchers in Hobart, Perth and Brisbane can tap into, and the national appetite for FAIR-aligned practice keeps growing. Yet the everyday realities of tropical humidity in Darwin, the time-zone friction between Adelaide and international collaborators, and the patchwork of university repositories across the country mean that conference datasets still slip through the cracks more often than anyone would like. The approaches below are drawn from a mix of international guidance and local experience, and they assume that the people reading this have neither unlimited budgets nor a dedicated digital preservation team.

Why early planning prevents late regret

Most conference data is generated under time pressure. A volunteer or postgraduate student captures audio on a phone, downloads slides from a cloud drive, and emails everything to a colleague at the end of the day. Months later, no one can find a particular file, or it turns out the recording was overwritten. Embedding a preservation plan into the conference itself is the single most effective habit, because it forces decisions to be made before files accumulate. This means writing down a retention schedule, assigning someone the role of data steward, and agreeing on a destination repository before the first session begins.

Australian events have started to adopt this discipline more visibly. The Australasian Open Access Strategy Group has been pushing conference proceedings onto persistent identifiers, and several state libraries, including the State Library of Queensland, now accept deposits of grey literature alongside journal articles. None of this happens by accident. A steering committee that meets four weeks before the event, again the day after, and once more a quarter later, will catch problems that an afterthought cleanup never will. The cost of a few meetings is dwarfed by the reputational benefit of a complete, well-described archive.

Choosing file formats that survive the decades

Format choice is where many preservation efforts quietly fail. A CSV opened in twenty years' time will probably still look like a CSV, but a proprietary project file from a discontinued statistics package will not. The general rule is to prefer open, documented, widely implemented formats and to keep the originals alongside any converted copies. For tabular data, comma-separated values, tab-separated values, and Parquet are all reasonable. For images, uncompressed TIFF and PNG remain safer choices than high-efficiency formats whose specifications may shift. Audio and video deserve extra care because codecs come and go, and migration to a current mainstream format every few years is healthier than a single conversion at the end of the project.

A quick comparison gives a starting point that can be adapted to most academic events.

File type Recommended formats Risk level Migration notes
Tabular data CSV, TSV, Parquet Low Validate delimiters; keep header rows intact
Spreadsheets ODS, XLSX Low to medium Avoid macros; export a CSV copy as well
Documents PDF/A-2, ODT, DOCX Low Prefer PDF/A-2 for the archival master
Images TIFF, PNG, JPEG2000 Low Keep uncompressed masters; use JPEG for derivatives
Audio WAV, FLAC Low Migrate to FLAC to save space; document sample rate
Video Matroska (FFV1), MP4 (H.264) Medium to high Plan transcoding every five to seven years
Presentations PDF/A, PPTX with embedded media Medium Export a flat PDF/A copy of every deck

Conference organisers who are not technical specialists often inherit files in whatever format presenters happen to send. A simple intake checklist helps. Ask for slides as PDF, raw data as CSV or ODS, audio as WAV or FLAC, and video as MP4 with the H.264 codec. Anything else triggers a quick conversation. Over time, this kind of polite gatekeeping shapes what people submit, and the archive becomes easier to maintain without anyone feeling blocked.

Storage architecture, backups and geographic separation

Storage strategy for a small archive is different from storage for a national collection, but the principles overlap. The classic three-copy, two-media, one-offsite rule still holds. One copy lives in the working repository where researchers can browse it. A second copy sits on a separate system, ideally in a different physical building, ready to be restored if the first is corrupted. A third copy, often called the dark archive, is held in a geographically distant location and is not accessed unless the other two fail. Australian institutions have an advantage here because the country's two main academic networks, AARNet and the various state government data centres, sit far enough apart to make genuine geographic separation straightforward.

The practical temptation, especially for small conferences run on a shoestring, is to keep everything in a single Dropbox folder, a Google Drive account, or a USB drive on someone's desk. Each of these has failed publicly and repeatedly. A cloud drive is fine as a working space, but it is not preservation. A USB drive is fine as a courier, but not as an archive. A minimal honest approach costs very little: a university research office will often provide a small allocation on its institutional repository, and pairing that with a second copy held by the conference's host library meets the basic two-copy rule without anyone needing to buy hardware.

Bit rot is the silent failure mode that catches people out. Storage media degrade, and checksums catch the problem before it becomes irreversible. Generating a SHA-256 checksum for every file on deposit, storing those checksums in a small manifest, and re-validating the manifest annually is a discipline that costs an afternoon to set up and a morning a year to maintain. This is the same approach recommended in the Data in the middle: practical examples from the conference, where organisers have documented the practical steps taken to safeguard their own materials.

Documentation, metadata and identifiers

A dataset without documentation is a sealed box. Future users may want to cite it, replicate an analysis, or simply understand what a particular column means. The FAIR principles, which Australian funders have endorsed through the ARDC's national statement, push researchers toward findability, accessibility, interoperability and reusability, and each of those depends on metadata. A good conference dataset will carry a title, author list, abstract, date, geographic scope, list of related publications, and a brief methods note. It will also have a persistent identifier, ideally a DOI minted through DataCite or a similar service, so that the citation never breaks when servers move.

The narrative around a dataset matters as much as the technical description. A two-page overview that explains why the conference was held, who funded it, what each session tried to achieve, and how the files fit together, turns a folder of orphaned objects into a coherent record. The same instinct shows up in fields that think explicitly about storytelling, and building a project narrative follows principles that translate surprisingly well into research contexts, where readers benefit from being told what to look at first. Good documentation answers four questions in plain language: what is this, who made it, what is missing, and what should I read next.

Discipline-specific metadata standards save enormous time when they exist. Social scientists reach for DDI, earth scientists for ISO 19115, librarians for Dublin Core, and biomedical researchers for CDISC. Conference datasets rarely fit a single standard, but borrowing a template from the closest one is a reasonable starting point. Australian researchers can also draw on the vocabularies maintained by the Australian Curriculum, Assessment and Reporting Authority for educational metadata, and by the National Archives of Australia for descriptive practice, both of which are freely available and well documented.

Long-term access, withdrawal and community trust

Preservation is not the same as perpetual publication, and an honest archive acknowledges that. Some datasets contain personal information and must be withdrawn or restricted after a set period. Others are superseded by better versions and should be marked as historical. A withdrawal notice, a tombstone page that explains why a file is no longer available and where to find the successor, respects the citation chain better than silent disappearance. Australian privacy law, particularly the Privacy Act and the notifiable data breaches scheme, sets a floor for what can be held openly, and conference organisers working with health, education or location data should consult their institutional research office before opening access.

Trust is the asset that preservation protects. Researchers who know that a conference dataset will be safely stored, clearly described, and stably cited are more willing to share work in progress, more willing to invite scrutiny, and more willing to cite their peers' contributions. Building that trust takes years and can be lost in a single careless incident. The conferences that succeed in this area tend to be the ones that publish a preservation policy on their website, list a contact email that actually works, and treat feedback as a gift rather than a complaint. The Australian Open Access Support Group and the Australasian Preservers of Data all offer informal advice to organisers who want to do this well and who would rather learn from someone else's mistake.

The ultimate test of a good archive is whether it can still be used in twenty years. That is a longer horizon than most conference budgets cover, so the practical answer is to deposit the materials in an institution that thinks in those terms. Universities, national libraries, and subject repositories such as those run by the Australian Ocean Data Network are built to outlast the events that fed them. Handing the archive over, with documentation, identifiers, and checksums in place, is the closing gesture of a well-run conference. Anything less, and the work of the participants quietly leaks away.

At the Conference

What attendees experienced in Lawrence

p1 reed

Plenary Sessions

Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.

platinum icpsr

Workshops

Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.

gold rockhurst

Social Events

An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.

Venue & Accommodations

Where the conference took place

silver ifdo

Kansas Union

University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.

silver ciser

The Eldridge

701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.

platinum icpsr

Springhill & TownePlace Suites

Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.

Program Highlights

Sessions and activities

Plenary Sessions Workshops Poster Presentations Committee Meetings Opening Reception Banquet Tour Kansas City Pecha Kucha Check-In Local Favorites

Getting Here

Lawrence, Kansas

Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045

Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).

Plan Your Stay

Accommodation options that were available

silver ifdo

The Eldridge

701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.

platinum ddi

The Oread

1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.

silver ciser

Springhill Suites

Marriott property. Room block reserved under "KU IASSIST Conference."

gold rockhurst

TownePlace Suites

Marriott property. Room block reserved under "KU IASSIST Conference."