Back

Scientific Software & AI Engineering for R&D Teams

I build reliable software at the boundary between models, scientific data and experiments: the data infrastructure, the evaluation, the AI workflows and the experimental tooling that turn a research capability into a system a team can depend on.

What I build

Most projects need several of these at once:

  • Scientific products and platforms: Python services, APIs and internal research tools that take an ambiguous research need to something scientists can use without its author beside them.
  • Scientific data infrastructure: ingestion, validation, structured metadata, provenance and versioned transformations, so a processed result can be traced back to the raw measurement and re-derived.
  • Evaluation and reliability: harnesses, regression suites, deterministic checks and documented failure modes, so a probabilistic system becomes testable rather than merely convincing.
  • AI workflows grounded in scientific data: agents, tool calling, structured outputs and retrieval over primary sources, with humans approving the steps that are expensive to undo.
  • Experimental design and decision systems: the layer that turns model output into the next experiment, choosing what to measure under a real budget.

From experiments to software

I started on the experimental side. My MSc thesis covered two years of wearable biosensor work at the Institute for Future Technologies: the BioWatch, a smartwatch I built from scratch to carry the enzymatic biosensors I was developing, and a microneedle lactate module for it. I then joined R&D at PKvitality as a research assistant, running in vitro tests on electrochemical microneedle CGM prototypes and documenting the anomalies that guided the next iteration.

Which is why I don’t treat scientific data as clean input. Measurements are noisy, instruments drift, protocols change between runs. It is the question I now ask of software: what happens when the data is wrong, and how would we find out?

Selected evidence

  • epibudget: an open-source tool that spends a fixed experimental budget on the protein variants exposing interaction structure, not the ones a model predicts will score well. Success criteria were registered before the results, and an earlier interpretation was withdrawn once an audit contradicted it.
  • Scientific Claim Verifier: an open-source engine checking each cited claim against the source it points to. Everything that does not need a model stays deterministic, every step emits provenance, and a regression guard fails the build if SciFact F1 drops below its committed baseline (0.92 against 0.62 naive, verifier-only).
  • Founder-level ownership: Finexov, an AI platform for public-funding applications taken from cold calls to €30K in sales, and Oseille AI, an agent for French innovation subsidies. Unclear problems, systems defined from scratch, nobody else to hand the ambiguity to.

How I think about AI for Science

The question I keep returning to is not what a scientific system can generate, but how it finds out that it is wrong. Feedback from reality has both a cost and a fidelity, and the two do not move together: the cheapest loops are the easiest to scale and the easiest to fool yourself with. Closing that gap, then using the result to choose the next experiment, is where I find the interesting engineering.

I wrote that argument out in AI for Science Is Moving From Prediction to Closed-Loop Research Systems, applied it to protein experiments in Measure for Information, Not for Fitness, and to research agents in Science Is Entering Its Agentic Era.

Who I work with

R&D teams at the point where a model, a prototype or a research workflow has to become dependable software: AI-native biology and TechBio startups, scientific platform teams, and AI-for-science groups inside larger organisations. Biology is where I am most fluent, but the work generalises wherever models, data and experiments have to line up.

I like working inside a team rather than beside it, close to the scientists who will use the system, where it is easier to see which part of their workflow actually breaks. I take one problem at a time and stay with it through the parts nobody could specify at the start, and I write production code with tests and explicit failure modes. Remote on CET hours with comfortable overlap for EU and US-East teams, and glad to relocate for long-term work.

Common questions

What stack do you work in?

Python: async services with FastAPI and Pydantic, typed and tested (mypy --strict, pytest). LLM APIs and MCP servers, retrieval over CrossRef, OpenAlex and PubMed, ESM-2 where the science calls for it, and Docker deployment, including the air-gapped environments I build under at LocusLab.

How do engagements start?

A 30-minute intro call. If the problem is a fit, I send a concrete proposal within a few days: scope, deliverables, timeline. If it isn’t, I’ll say so directly.

Work with me

I take on selective freelance engagements with biology, TechBio and AI-for-science R&D teams: scientific data infrastructure, evaluation and reliability, AI workflows, and the systems that turn model output into the next experiment.