Two complementary workflows
A closed intervention dossier first needs to be brought together, assessed and documented. It can then support two related purposes: keeping a trustworthy long-term record and sharing appropriate research data.
Select, organise & document
Paper, hybrid and digital-born intervention dossiers
Archive & safeguard
Archival packages, integrity and long-term storage
Publish & discover
FAIR metadata, repository access and search
The workflows are evaluated with the same representative test set, so problems found in practice can inform the model and implementation.
Building on shared infrastructure
The pre-archive and MetaHub
HESCIDA introduced a central pre-archive for organising closed digital dossiers and MetaHub for documenting their datasets. H-SEARCH proposes adapting this approach to historical and hybrid records.
BALaT
The institutional knowledge platform connects object descriptions, photographs, library records and research data. The project proposes extending discovery from metadata to the contents of scientific documents.
Dataverse
The linked repository provides access to datasets described through the knowledge platform. Persistent identifiers support stable references to dossiers and their data.
Elasticsearch
The search infrastructure aggregates information from different sources. H-SEARCH proposes adding multilingual and full-text capabilities to support richer retrieval.
Archive Wrapper: an implemented project tool
Developed for H-SEARCH, Archive Wrapper provides a guided interface over the Archivematica pipeline. It connects selected dossier files to granular metadata from the CORDRA FAIR Digital Objects repository and follows their progress into Archival Information Packages (AIPs). A separate RO-Crate mockup explores a possible extension.
Long-term preservation infrastructure
The preservation workflow is designed to use BELSPO’s LTP platform and KIK/IRPA’s tape infrastructure. Established archival standards and packaging solutions are to be assessed during WP.2.
Finding knowledge inside documents
For a scanned report, optical character recognition (OCR) is a first step toward machine-readable text. Digital-born documents can supply text more directly. Both streams then need to be integrated into a consistent discovery workflow.
The proposal investigates natural language processing, word embeddings and specialised thesauri to expand queries across languages. It also evaluates automated summaries and translation to help users understand retrieved material.
A simple example of the intended experience
A researcher looking for information about stone sculpture should be able to discover relevant French or Dutch reports as well as English ones. This illustrates the research ambition; this website does not provide a live archive search.
Responsible access
Scientific documents can sit alongside personal details, correspondence and administrative information. The project plans guidelines for distinguishing research data from material requiring restrictions, including legal review of GDPR-related guidance.
FAIR access therefore includes clear conditions for use and appropriate access levels, alongside rich metadata, persistent identifiers and documented context.
Read about data management →