Dries Huybens

August 13, 2026 Daniel Hladek erasmus 1 minute, 14 seconds

ERASMUS Intern Summer 2026, 20 July - 30 August

Topic:

Multilingual Knowledge Graphs from medical data

Goal:

  • Construct a knowledge graph from medical package inserts in multiple languages
  • Utilize the graph in an intelligent agent that recommends medication.

Tasks:

  • Continue project of Bogdan Paul Chis
  • Study repositories:

    • https://github.com/chis-facultate/erasmus-kosice
    • https://github.com/hladek/mul-me-kg
    • https://github.com/hkuds/lightrag
  • Learn intelligent agents and generative models - OpenAI API, Agent frameworks, RAG systems.
  • Learn about knowledge graphs and GraphRAG. Read several research papers.
  • Prepare a Python based workflow, use git code repository
  • Visualize the graph
  • Prepare an agent that utilizes the unstructured data and graph-data.
  • Evaluate the agent using DeepEval or RAGAS.
  • Write a report
  • Put all code to GIT

Project tasks update:

  • Prepare a multilingual parallel corpus from OPUS data
  • Prepare dataset for visual language model evaluation from foto
  • Prepare dataset for visual language model evaluation from wikipedia. You can use https://huggingface.co/datasets/wikimedia/wit_base .
  • Prepare corpus of text and image data from pravda.sk . For each subdomain, create - raw HTML data with images. Then create corpus of images plus descriptions, subtitles and tags. Then create corpus of extracted text. You can use docling , trafilatura for text extration. Or design your own parser. Give source codes to a git repository with documentation. Put data files to school server. Do not download too fast, so our IP does not reveive a ban.

Outupus:

  • multilingual corpus was analyzed with embedding models. Semantic overlap between languages is low, so multilngual parallel corpus will be small.
  • A corpus from FotkyZadarmo.sk