Turning Conference Records Into Reusable Research Data
Conference materials often look like finished communication rather than research inputs. A programme PDF, presentation deck, plenary recording or delegate guide may have been created for a few days in one place, then left online as an archive. Yet these records can reveal research themes, institutional relationships, emerging methods and the way a field understood its own priorities at a particular moment.
The IASSIST 2017 archive is a useful example. Held at the University of Kansas in Lawrence from 23 to 26 May 2017, “Data in the Middle: The Common Language of Research” brought together work on big data, deep learning, digital agriculture and global food security. Its schedule, plenary information and practical attendee material form a small but valuable research collection.
Learning how to reuse conference data means treating these materials as evidence with provenance, rather than copying isolated facts into a spreadsheet. The work involves discovery, rights checking, structured extraction, cleaning, interpretation and careful citation. It can support a literature review, a historical study of research communities, a dataset about conference topics or a teaching activity in a university course.
For Australian researchers, the same approach applies to records from a uni seminar, a CSIRO workshop, an industry expo or a professional meeting in Sydney, Brisbane, Canberra or a regional centre. A useful workflow makes old event material searchable, comparable and fit for a new purpose without stripping away the context that gives it meaning.
Define The Research Use Before Extracting Data
Begin by stating what you want the conference material to help you discover. A broad aim such as “analyse the conference” is too vague to guide consistent collection. A sharper question might examine how often machine learning appeared alongside food security, which disciplines were represented in plenary sessions, or how digital agriculture was described in 2017 compared with current Australian policy language.
Your intended output determines the fields you need. A topic analysis may require session titles, abstracts, speakers, affiliations, keywords and presentation types. A network study may focus on names, institutions, countries and collaborations. A study of scholarly communication might collect slides, videos, links, references and evidence of audience engagement. Keep a short data dictionary that explains what each field means and how it will be recorded.
It is also worth separating primary event records from your interpretation. The published programme is evidence of what organisers scheduled, while your labels for “deep learning”, “agriculture” or “interdisciplinary” are analytical decisions. Store both layers. This distinction becomes important when another researcher checks your coding or when you later revise the categories.
A practical scope prevents the project from becoming an unmanageable scrape of every page. For the IASSIST archive, you might start with plenaries and research presentations, then add practical pages only if they help explain attendance, location or institutional context. In Australia, a similar project could begin with one annual meeting of a professional association rather than attempting to process decades of conference websites.
Locate, Capture And Preserve The Source Material
Archived websites rarely behave like modern databases. Information may be distributed across navigation pages, downloadable files, embedded media and links to external platforms. Record the page title, original URL, access date and file type as soon as you find an item. Save a stable local copy where the terms of use permit it, and retain the surrounding page so that a presentation file does not become detached from its session or author.
A source inventory can include the item identifier, session, speaker, organisation, format, publication date, language, rights statement and preservation status. Give files consistent names, such as 2017_plenary_speaker_topic.pdf, while keeping the original filename in a separate field. Generate checksums for important files if the project requires strong evidence that a downloaded copy has not changed.
Do not assume a PDF is a dataset. It may contain selectable text, scanned pages, photographs of whiteboards or charts whose values cannot be read reliably by software. Optical character recognition can help with scans, but every extracted passage needs quality checking. Presentations may also contain duplicated title slides, missing fonts, speaker notes or references that point to resources no longer online.
Media deserves its own record. If a conference recording exists, capture its title, duration, presenter, URL, transcript status and any access restrictions. When tracing audiovisual files or production credits, an external media provider such as B-side Productions may offer useful context about how event content was created and delivered. Treat the recording, transcript and production information as related but separate objects.
Check Rights, Consent And Cultural Context
Reuse begins with permission, not extraction. A conference organiser may have permission to display a slide deck without having permission to redistribute every photograph, dataset, logo or recording within it. Look for copyright notices, Creative Commons licences, speaker agreements and terms applying to the archive. If the status is unclear, document the uncertainty and contact the rights holder before publishing a copy.
A rights register can use simple categories such as open to reuse, available for analysis only, permission required, or do not redistribute. You may be able to quote a title and cite a page while being unable to republish the entire PDF. For recordings, consider voice, image and performance rights. A transcript can also contain identifiable statements that require careful handling, especially if the material was not created for secondary research.
Australian projects need to consider local expectations around Indigenous data. The CARE Principles—Collective Benefit, Authority to Control, Responsibility and Ethics—add considerations that are not covered by technical openness alone. A conference presentation involving Aboriginal or Torres Strait Islander communities, Country, language or cultural knowledge may require community guidance before its content is classified, combined or visualised.
The same principle applies to agricultural and commercial data. A presentation about farm operations, crop yields or supply chains may include confidential information even if the slides remain publicly viewable. Separate “available online” from “safe to republish”. A sensible approach is to publish metadata and an access pathway while keeping sensitive files under controlled conditions.
Clean And Enrich The Extracted Dataset
Conference records are inconsistent by nature. One speaker may appear as “Dr Jane Smith”, “Jane Smith” and “J. Smith”; one institution may be listed by a full legal name and elsewhere by an acronym. Normalise names and organisations in separate fields rather than overwriting the wording used in the original source. Preserve the raw value, add a cleaned value and record the rule used to make the change.
Dates need similar care. Store the event date, presentation date and publication or upload date separately. This matters when comparing an Australian conference held across an AEST boundary with an international event displayed in local Kansas time. A schedule can also contain cancellations, moved sessions and repeated items, so retain status notes rather than treating every listed presentation as delivered.
Controlled vocabularies make comparison possible, but they should not flatten meaning. You might create topic tags for data infrastructure, machine learning, agriculture, food systems, policy and research support, while allowing multiple tags per item. Keep the original title and abstract alongside the tags. If two people code the material, compare their decisions on a sample and discuss disagreements before processing the full collection.
Enrichment can add value when it is transparent. Link a speaker to an ORCID record where the match is certain, connect an institution to a persistent identifier, or add a DOI for a cited publication. Do not infer identity from a name alone. For Australian work, resources associated with the Australian Research Data Commons can help frame repository and metadata decisions, while institutional research offices may have local requirements for records deposited through a university system.
Analyse, Cite And Publish A Reusable Result
Once the collection is structured, choose analysis methods that match its limits. Frequency counts can show which subjects dominated the programme, while co-occurrence analysis can identify topics appearing together. A timeline may reveal when terms such as “deep learning” moved from specialist sessions into wider research discussions. Qualitative coding is often better for interpreting how speakers defined “big data” or “global food security”.
Use the archive’s own structure as evidence. The order of sessions, the prominence of plenaries and the wording of event themes all carry meaning. An article about interdisciplinary research can help situate the conference’s emphasis on data as a shared language, but it should supplement rather than replace your record of the programme and presentations.
A reusable release should include a readme, data dictionary, source list, licence information, processing notes and a version number. Explain which records were unavailable, how duplicates were handled, which names were disambiguated and whether machine transcription was reviewed. Deposit the structured dataset in a suitable repository with a persistent identifier where possible. If full files cannot be shared, publish a metadata-only version with instructions for responsible access.
Citations should make the path back to the evidence clear. Cite the event, session, presenter, item title, archive URL and access date when a persistent identifier is unavailable. For derived statistics, cite the dataset version and describe the transformation. A reader should be able to distinguish a claim made by a speaker in 2017 from an observation produced by your later analysis.
| Reuse Project | Core Material | Useful Fields | Main Risk | Suitable Output |
|---|---|---|---|---|
| Topic history | Programme and abstracts | Title, date, session, keywords, abstract | Categories impose present-day meanings | Coded dataset and timeline |
| Research network | Speaker and affiliation records | Name, role, organisation, country, links | Names may refer to different people | Network map with provenance |
| Presentation analysis | Slides, transcripts and recordings | File type, transcript, references, rights | Copyright and transcription errors | Searchable research corpus |
| Digital agriculture study | Relevant sessions and cited sources | Crop, region, method, institution, year | Commercial or farm information may be sensitive | Thematic review and metadata record |
| Teaching resource | Selected archive items | Learning objective, source, licence, activity | Students may confuse archive with current evidence | Curated classroom collection |
A clear preservation routine is especially valuable in Australia, where researchers often move between university repositories, government portals and industry systems. A dataset prepared for an ARC-funded project may need a different access model from a public teaching collection. Keeping the source, method and permission trail together means the work remains useful when staff change, a website is redesigned or an old conference becomes relevant to a new question.
The strongest reuse does not pretend that an event archive is complete. It acknowledges missing presenters, inaccessible files, inconsistent metadata and the historical limits of the programme. That honesty makes the result more credible. A carefully documented collection can turn a few days of conference activity into evidence for research history, information studies, policy analysis and the changing practice of data-intensive scholarship.
At the Conference
What attendees experienced in Lawrence
Plenary Sessions
Keynotes from Daniel Reed on data, technology, and culture, and Jennifer Clarke on digital agriculture and the Midwest Big Data Hub.
Workshops
Full-day technical workshops on Tuesday, May 23. Morning sessions ran 9:00–12:00 and afternoon sessions 13:00–16:00. Laptops were required.
Social Events
An opening reception at The Oread, a banquet, and an optional post-conference tour of Kansas City including Crescent Moon Winery.
Venue & Accommodations
Where the conference took place
Kansas Union
University of Kansas campus, Lawrence. Main conference venue with check-in on the 4th and 5th floor lobbies.
The Oread
1200 Oread Avenue, Lawrence. Hosted the opening reception and offered a room block for attendees.
The Eldridge
701 Massachusetts Street, Lawrence. A partner hotel with a reserved room block for conference guests.
Springhill & TownePlace Suites
Marriott properties in Lawrence with room blocks reserved under the "KU IASSIST Conference" name.
Program Highlights
Sessions and activities
Getting Here
Lawrence, Kansas
Kansas Union · University of Kansas
1301 Jayhawk Blvd, Lawrence, KS 66045
Kansas City International Airport (MCI) is approximately 50 minutes by car. Lawrence Transit Routes 10 and 11 served the area ($1 exact change).
Plan Your Stay
Accommodation options that were available
The Eldridge
701 Massachusetts Street, Lawrence, KS 66044. Room block now closed.
The Oread
1200 Oread Avenue, Lawrence, KS 66044. Room block now closed.
Springhill Suites
Marriott property. Room block reserved under "KU IASSIST Conference."
TownePlace Suites
Marriott property. Room block reserved under "KU IASSIST Conference."