Purpose and approach
A metadata search can identify a dossier without revealing the knowledge inside it. H-SEARCH proposes searching the documents themselves, using OCR for scanned text and language technologies to help bridge Dutch, French, English and other languages present in the archive.
Tasks & planned deliverables
Month numbers are relative to the proposed project start. They do not indicate confirmed completion dates.
Task 3.1 · Months 7–18
Model the dissemination workflow
Adapt the dissemination model developed in HESCIDA for paper, hybrid and digital-born records. Build on BALaT and Dataverse, with appropriate metadata and access levels.
Task lead: Erik Buelinckx
Planned deliverables
- D.3.1.1State of the art and draft dissemination workflowMonth 12
- D.3.1.2Workflow model for dissemination of research dataMonth 18
Task 3.2 · Months 7–30
Develop multilingual full-text search
Investigate specialised thesauri and natural language processing word embeddings for multilingual query expansion. Implement search in Elasticsearch and BALaT, and evaluate automated summarisation and machine translation for retrieval.
Task lead: Roald Hayen
Planned deliverables
- D.3.2.1Guideline for integrating word embeddings and specialised thesauriMonth 15
- D.3.2.2Multilingual search platform onlineMonth 27
- D.3.2.3Performance report on automated machine translationMonth 30
Task 3.3 · Months 15–30
Evaluate the repository workflow
Use the WP.1 test set to evaluate the workflow and its implementation. Return findings to the modelling and development tasks to improve the resulting service.
Task lead: Véronique Van der Stede
Planned deliverables
- D.3.3.1Evaluation reportMonth 30
How this work connects
This work extends HESCIDA’s infrastructure and uses the same test dossiers as WP.2. The aim is to link discoverability, contextual information and access to the underlying data.