Archiving Conference Data: Lessons from IASSIST 2017
A conference website can look temporary while it is active, yet become a valuable research record once the event has finished. The IASSIST 2017 website demonstrates this clearly. Created for “Data in the Middle: The Common Language of Research”, held at the University of Kansas in Lawrence from 23 to 26 May 2017, it preserves far more than a timetable. It records the themes, people, presentations and practical decisions that shaped a professional gathering.
For archivists, research data specialists and university libraries, the site offers a useful case study in archiving conference data. Its programme includes sessions on big data, deep learning, digital agriculture and global food security, while its practical pages cover check-in, transport and local recommendations. Together, these materials show how an event archive can capture both intellectual content and the lived experience of attendance.
Australian institutions face similar preservation questions. A conference hosted in Sydney, Melbourne, Brisbane or Perth may generate registration records, streamed sessions, speaker biographies, slide decks, photographs, social media posts and venue information. Without a planned approach, these materials can disappear when a content management system is retired, a domain expires or a cloud subscription changes.
Defining The Conference Record
The first step is to decide what belongs in the archival collection. The obvious items are the official programme, schedule, abstracts, speaker profiles, plenary details and presentation files. These establish what was planned and who contributed. However, the record can be weakened if it excludes registration guidance, venue maps, travel information or announcements that explain how the event operated.
IASSIST 2017 illustrates the value of preserving context. A session title such as one relating to deep learning or digital agriculture provides a broad subject signal, but the associated presenter, abstract, time slot and presentation file create a much richer record. A future researcher can see how the field was discussed in 2017, which concepts were prominent and how those concepts were positioned within the wider research data profession.
The collection should be described as a set of related objects rather than a single website snapshot. A schedule may link to a plenary page, which links to a presentation and an author biography. Recording these relationships helps users navigate the archive and allows search systems to expose meaningful connections between people, topics and sessions.
Capturing Web Content Before It Changes
Conference websites are dynamic. A page may display content from a database, generate links through scripts or draw images from a separate server. A basic download of visible HTML may therefore miss documents, menus, registration notices and media files. Web harvesting should capture the site at more than one level, combining a crawl of public pages with a carefully assembled collection of linked files.
A preservation copy should retain the original URLs, retrieval dates and technical response information. These details help distinguish a page that was available during the event from one that was added later. They also support authenticity checks if a file is challenged or if the archive is migrated to another platform.
Screenshots and PDF renderings have a place, particularly for pages whose layout carries meaning. For example, a visual schedule may show parallel sessions, room changes and plenary placement more clearly than extracted text. Still, rendered copies should supplement structured web content rather than replace it. HTML, plain text and machine-readable metadata are easier to search and reuse over time.
For Australian universities, a practical workflow may involve a web archiving service alongside the institution’s repository. A conference hosted in Melbourne could retain a WARC package for the website, while the library repository stores presentation files and descriptive records. This separation reduces dependence on the original event platform and provides a second preservation pathway.
Preserving Presentations And Research Materials
Presentation files often form the intellectual core of a conference archive. PowerPoint files may contain embedded video, linked fonts, speaker notes or external charts. PDFs are usually easier to preserve and distribute, but converting every file without retaining the original can remove useful evidence. The preferred approach is to keep the submitted file, create an access version where necessary and document any conversion.
File names should be normalised without erasing the original name. A consistent pattern might include the year, event acronym, session code, presenter surname and version status. Technical metadata should record the file format, size, checksum, creation date and preservation action. Fixity checks using a cryptographic hash can reveal whether a file has changed after ingest.
The same discipline applies to recordings, photographs and supplementary datasets. Audio and video should be stored in stable, well-documented formats, with separate captions or transcripts where available. Photographs require descriptions, dates and rights information. A slide deck that refers to a dataset should ideally link to the dataset’s persistent identifier rather than an unstable download address.
The themes represented at IASSIST 2017 also highlight the importance of contextual metadata. A presentation about global food security might refer to geographic regions, institutions, crops or statistical sources. A record containing only the file title will have limited discovery value. Subject keywords, abstracts, geographic coverage and related project names make the material useful to researchers well beyond the original audience.
Recording People, Rights And Consent
Conference data frequently contains personal information. Registration forms may include names, affiliations, dietary requirements, accessibility needs and contact details. Speaker biographies may include email addresses and photographs. Attendance lists, chat logs and photographs can create additional privacy risks. These materials should be divided into public, restricted and confidential categories before they enter a long-term archive.
Australian organisations need to consider the Privacy Act 1988, applicable state or territory requirements and their own institutional policies. A public university in Queensland may apply different operational procedures from a private event organiser in New South Wales, yet both need a clear legal basis for retaining and disclosing personal information. Sensitive registration details should normally be excluded from the public archive or stored under controlled access.
Copyright and licensing decisions should be recorded at the point of collection. A presenter may permit a slide deck to be available to registered attendees but not to the general public. Another may use third-party images, licensed datasets or publisher-owned material that prevents open redistribution. A rights statement should identify the copyright holder, permitted uses, access limits and any required attribution.
Consent forms should cover recording, online publication, reuse and future preservation rather than only the live event. If a plenary is filmed, speakers should understand where the recording will be hosted and for how long. In Australian practice, an acknowledgement of Country may be part of the opening ceremony and should be preserved with the event record, while avoiding the assumption that it grants permission to reproduce every associated image or statement.
Designing Metadata For Discovery
A dependable archive needs a metadata profile that is detailed enough to support discovery without becoming impossible to maintain. Core fields usually include title, creator, contributor, date, description, subject, event name, session type, location, language, file format, rights and persistent identifier. Controlled vocabularies should be used for recurring values such as “plenary”, “paper”, “poster” and “workshop”.
Event metadata should distinguish between the conference as a whole and each programme item. A session record can contain its start and end times, room, chair, presenters, abstract and related files. Time zones and local dates matter when an event has international participation. Although IASSIST 2017 was held in Lawrence, an Australian archive should state the local time zone explicitly for online or hybrid sessions, especially when attendees join from Adelaide, Darwin or Perth.
Persistent identifiers add stability. An institution may assign a DOI to the complete collection and repository identifiers to individual presentations. ORCID iDs can distinguish researchers with similar names, while organisational identifiers can separate universities with comparable titles. Links should be tested regularly because an archive full of broken references quickly loses practical value.
The Australian Research Data Commons and university repository networks provide useful reference points for local data stewardship. Archives can align their descriptive fields with institutional research data catalogues, making conference material discoverable beside datasets, reports and project outputs. This is especially valuable when a conference presentation is an early public account of research that later appears in a formal publication.
Choosing Storage, Access And Preservation Formats
Storage should be planned in layers. The preservation master belongs in managed institutional storage with access controls, backups and routine integrity checks. A separate access copy can be optimised for web delivery, while a disaster recovery copy is held in another location or service. Cloud storage may be convenient, but procurement should examine data residency, exit arrangements, service-level commitments and the cost of retrieving large collections.
Format selection should favour open or widely supported standards. Plain text, XML, CSV, PDF/A, TIFF, WAV and Matroska are common preservation candidates, although the correct choice depends on the material. Proprietary formats should not automatically be rejected; retaining the original can preserve features that conversion removes. The key is to document dependencies and create a usable access version.
A preservation plan should identify who owns the archive after the organising committee dissolves. Universities often have a natural home in their library or research office, while professional associations may need a consortium arrangement. An event held in Adelaide, for example, could be hosted by a national association but preserved through the university’s repository under a written agreement defining costs, rights and responsibilities.
Costs should include migration, metadata creation, access support and periodic review, not just initial storage. Australian organisations may compare local data centre services with international cloud platforms, taking account of the Australian dollar, GST, public-sector procurement rules and the value of local technical support. A cheaper storage tier is not economical if it makes retrieval slow or incurs unexpected egress charges.
Comparing Preservation Approaches
There is no single archive model suitable for every conference. A small workshop may need a curated repository deposit, while a large international event may require a complete web crawl, audiovisual preservation and a rights management process. The appropriate choice depends on the event’s scale, the sensitivity of its records and the expected research value.
The following comparison separates common approaches by what they preserve, where they work well and what risks require attention:
| Approach | Materials preserved | Strengths | Main risks |
|---|---|---|---|
| Basic web snapshot | Public pages, navigation and visible media | Fast, affordable and useful for documenting the public site | Linked files, databases and restricted content may be missed |
| Curated repository deposit | Programme, abstracts, presentations and metadata | Strong discovery, stable identifiers and institutional stewardship | Requires selection, rights review and staff time |
| Full web archive with WARC files | Website structure, pages, scripts and linked resources | Closest representation of the online event experience | Dynamic content, external services and embedded media may not replay correctly |
| Multimedia preservation package | Video, audio, captions, photographs and transcripts | Retains the event’s spoken and visual record | Large storage needs, consent issues and format migration |
| Hybrid institutional archive | Web capture, repository files and controlled records | Balances authenticity, discovery and privacy | Requires coordination between organisers, IT staff and archivists |
IASSIST 2017 is particularly suitable for a hybrid model. Its archived public pages provide the event framework, while programme records and presentations can receive richer metadata in a repository. Practical attendee information can be retained as historical context, whereas private registration data should remain restricted or be securely disposed of according to policy.
Making The Archive Durable
Long-term preservation is an ongoing service rather than a final upload. Every few years, staff should test links, validate checksums, review file formats and confirm that the archive can still be searched. A short preservation log should record migrations, software changes, access restrictions and decisions about missing material.
User testing is equally important. A researcher should be able to locate a session from a topic, speaker or date, open the associated abstract and understand which files are available. A student examining digital agriculture should not need to know the original website’s menu structure. Clear landing pages, descriptive filenames and accessible transcripts make the collection useful to audiences beyond the conference community.
Archives should preserve evidence of uncertainty. If a presentation is missing, if a recording has incomplete captions or if a page could not be harvested, the record should say so. Transparent gaps are more trustworthy than silent omissions. Version histories are also valuable where speakers revised slides after the event or where a programme changed during the four-day schedule.
A durable conference archive ultimately connects event administration with scholarly communication. The IASSIST 2017 materials show how a professional meeting can remain valuable years later when its programme, ideas, people and setting are described together. For Australian libraries, associations and universities, the same principles support reliable digital stewardship across changing platforms, funding cycles and research priorities.
At the Conference
What attendees experienced in Lawrence
Plenary Sessions
Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.
Workshops
Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.
Social Events
An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.
Venue & Accommodations
Where the conference took place
Kansas Union
University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.
The Oread
1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.
The Eldridge
701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.
Springhill & TownePlace Suites
Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.
Program Highlights
Sessions and activities
Getting Here
Lawrence, Kansas
Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045
Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).
Plan Your Stay
Accommodation options that were available
The Eldridge
701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.
The Oread
1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.
Springhill Suites
Marriott property. Room block reserved under "KU IASSIST Conference."
TownePlace Suites
Marriott property. Room block reserved under "KU IASSIST Conference."