i'm emilia,
and i do research on
multi-agent llms.
getting them to play nice, figuring out what happens when they don't, and how anyone is supposed to check their work.
about
hi, i'm emilia. i'm a phd student in the ellis program, split between ellis alicante and the university of toronto, advised by nuria oliver and zhijing jin.
my thesis is about verifying multi-agent work: when a pile of agents hands you a result, how do you actually know they did the job right? humans get tired and miss things, model monitors have their own blind spots, and hybrid protocols sit somewhere in between. i build benchmarks that measure the verifier rather than the agent.
before the phd i worked as an ml engineer and researcher: llm safety and polish-language models at nask (including pllum, poland's first open-source llm), security ml at google, on-device nlp at samsung, and clinical llm systems at jutro medical.
publications
2025
Behind Closed Words: Creating and Investigating the forePLay Annotated Dataset for Polish Erotic Discourse
OpenGVL — Benchmarking Visual Temporal Progress for Data Curation
2024
BAN-PL: a Novel Polish Dataset of Banned Harmful and Offensive Content from Wykop.pl
Big Tech Influence over AI Research Revisited: Memetic Analysis of Attribution of Ideas to Affiliation
What Matters in Hierarchical Search for Combinatorial Reasoning Problems?
2023
Deep Dive into the Language of International Relations: NLP-based Analysis of UNESCO's Summary Records
Towards Harmful Erotic Content Detection through Coreference-Driven Contextual Analysis
2022
Climate Policy Tracker: Pipeline for Automated Analysis of Public Climate Policies
* equal contribution. full list also on google scholar.
cv
full pdf here. short version below.
education
- 2026 — 2029 ph.d., machine learning · ellis phd program — ellis alicante & university of toronto advisors: nuria oliver, zhijing jin · thesis: verifying multi-agent work — human cognitive limits and adaptive oversight in hybrid ai teams
- 2021 — 2024 m.s., machine learning · university of warsaw advisors: piotr sankowski, paweł budzianowski · grade: 5 · thesis later published at ACL 2025
- 2018 — 2021 b.s., mathematics · university of warsaw thesis: predicting escalations in customer support
experience
- 2026 — present research assistant · jinesis lab, university of toronto multi-agent llm systems · human oversight of agent collectives
- 2026 — present technical member · eurosafeai multi-agent ai safety · systemic risk research
- 2026 ml engineer · jutro medical pre-visit patient interviews (text & voice) · automatic clinical documentation with real-time evaluation · multi-agent chat over patient records with rag · e2e evaluation suite with blind a/b testing
- 2026 ai safety researcher · algoverse subliminal learning in chain-of-thought reasoning · mentor: michael mulet
- 2025 ml engineer · google (via dataphant) ml pipelines for security workflows: incident response, code review, legal compliance · deployed models estimated to save ~$6m/year
- 2024 nlp engineer · samsung lightweight on-device text features for mobile · literature surveys, paper authoring, peer review
- 2022 — 2024 senior nlp researcher · nask (national research institute) llm training & finetuning for polish and english (pllum) · llm safety: jailbreak resistance, refusal behaviour, harmful-content detection · csam detection in narrative text · supervised junior researchers
- 2022 — 2024 research software engineer · mi2datalab climate policy tracker · unesco tension detection · memetic analysis of big tech in ai research
- 2021 — 2022 python engineer · silver bullet solutions gpt-based chatbot & ner finetuning pipeline for a telecom client · anonymization and storage tooling for sensitive corpora
selected projects
- foundation model pllum · poland's first open-source llm, at nask pretraining & instruction data curation · pre-training and finetuning · systematic comparison of ppo / dpo / orpo · annotation guidelines for summarisation, preference and toxicity labelling
- benchmark opengvl · open source temporal task-progress prediction with vlms across robotic and human embodiments · used to curate large-scale robotics datasets
- benchmark adversarial multiple-choice robustness · m.sc. thesis sota llms degrade sharply when correct options are swapped for plausible-but-wrong distractors
- deployed system climate policy tracker · mi2datalab nlp pipeline classifying national climate policies, live as a public dashboard
- datasets harmful-content detection in polish · nask forePLay and ban-pl datasets · coreference-driven pipeline improving recall on long-form content
teaching & service
- 2019 — present organizer · ml in pl association project lead of the 2023 conference · current board member · long-time organizer across student research workshop, call for contributions, speakers, sponsorship & marketing teams
- 2023 — present member · unesco chair on intangible cultural heritage in public and global governance university of warsaw · faculty of political sciences and international studies
- 2022 — 2025 teaching assistant · university of warsaw intro to cs · intro to ml · deep neural networks · nlp · nlp for social scientists · weekly labs for groups of ~20
toolbox
- ml / research pytorch · transformers · langchain · smolagents · spacy llm pre-training, sft, ppo, dpo, orpo · distributed training with slurm
- programming python · c/c++ · r · rust · sql · git · docker
- languages polish (native) · english (fluent) · french & german (basic) plus scottish gaelic and spanish, in progress
misc / fun
the bits that don't fit anywhere else.
🐈 the cats
i have five. their names are Buba, Śmieciuch, Żepet, Kudłacz i Łapa.
📚 currently reading
- 1Q84 — Haruki Murakami
- Czekając na Godota — Samuel Beckett
- 80,000 Hours — Benjamin Todd
🎮 playing
the witcher 3.
🌱 learning
scottish gaelic & spanish.
say hi
best way to reach me: emilia@goral.one
i read most cold emails but reply slowly. if i miss yours, it's not personal — just nudge me again in a week.