Projects
LHCb Analysis Productions
Python · FastAPI · React · TypeScript · Celery · MySQL · SQLAlchemy · Kubernetes
Data collected by LHCb requires further processing to be useful for physics analysis, which depends on the kind of physics an analyst is looking for and/or trying to measure.

Not too long ago, an LHCb analyst would produce an analysis dataset by submitting their own jobs to the Worldwide LHC Computing Grid (WLCG) or simply the Grid (a collection of computing centres which execute jobs a bit like in a batch system), and effectively babysitting them: monitoring progress, diagnosing failures job by job, resubmitting, and preserving the outputs and configuration by hand. It was the most time-consuming and error-prone part of starting an analysis, and when LHC Run 3 multiplied the data rates this became completely unworkable.

Analysis Productions replaces that with a declarative, centrally managed service: the analyst describes what they need in a short YAML file, and the platform handles validation, CI testing (using real Grid jobs), submission, monitoring, and preservation, with provenance kept end to end so that any dataset can be traced back to exactly how it was made.

The platform grew out of LHCb's earlier working-group productions system and was already running when I joined the collaboration back in 2020. Since then, I have worked on much of the API backend and (completely new, overhauled) web frontend LbAPWeb (see below), and I remain one of the platform's core developers and maintainers.
It has been hugely successful. Analysis Productions is the official processing route for LHCb's Run 3 data, used by hundreds of analysts and open to the whole collaboration. As of August 2026 it had read 1.8 exabytes of collision data and returned 7.6 PB of analysis-ready datasets, a reduction of more than 200 to 1. Chris Burr presented the platform at CHEP 2026, as an exascale analysis data processing and management service for LHCb.
Data table
| Date | Samples created / month |
|---|---|
| 2020-01-01 | 42 |
| 2020-02-01 | 133 |
| 2020-03-01 | 188 |
| 2020-04-01 | 314 |
| 2020-05-01 | 100 |
| 2020-06-01 | 117 |
| 2020-07-01 | 134 |
| 2020-08-01 | 28 |
| 2020-09-01 | 46 |
| 2020-10-01 | 66 |
| 2020-11-01 | 88 |
| 2020-12-01 | 198 |
| 2021-01-01 | 286 |
| 2021-02-01 | 271 |
| 2021-03-01 | 258 |
| 2021-04-01 | 390 |
| 2021-05-01 | 546 |
| 2021-06-01 | 70 |
| 2021-07-01 | 1175 |
| 2021-08-01 | 126 |
| 2021-09-01 | 36 |
| 2021-10-01 | 314 |
| 2021-11-01 | 192 |
| 2021-12-01 | 96 |
| 2022-01-01 | 144 |
| 2022-02-01 | 245 |
| 2022-03-01 | 516 |
| 2022-04-01 | 206 |
| 2022-05-01 | 606 |
| 2022-06-01 | 525 |
| 2022-07-01 | 80 |
| 2022-08-01 | 226 |
| 2022-09-01 | 592 |
| 2022-10-01 | 372 |
| 2022-11-01 | 836 |
| 2022-12-01 | 158 |
| 2023-01-01 | 241 |
| 2023-02-01 | 444 |
| 2023-03-01 | 1400 |
| 2023-04-01 | 831 |
| 2023-05-01 | 363 |
| 2023-06-01 | 511 |
| 2023-07-01 | 393 |
| 2023-08-01 | 368 |
| 2023-09-01 | 1236 |
| 2023-10-01 | 1682 |
| 2023-11-01 | 381 |
| 2023-12-01 | 697 |
| 2024-01-01 | 720 |
| 2024-02-01 | 888 |
| 2024-03-01 | 792 |
| 2024-04-01 | 1359 |
| 2024-05-01 | 1141 |
| 2024-06-01 | 957 |
| 2024-07-01 | 1593 |
| 2024-08-01 | 1037 |
| 2024-09-01 | 1308 |
| 2024-10-01 | 1414 |
| 2024-11-01 | 2432 |
| 2024-12-01 | 1238 |
| 2025-01-01 | 1561 |
| 2025-02-01 | 3509 |
| 2025-03-01 | 1399 |
| 2025-04-01 | 939 |
| 2025-05-01 | 1776 |
| 2025-06-01 | 1535 |
| 2025-07-01 | 2806 |
| 2025-08-01 | 1705 |
| 2025-09-01 | 1779 |
| 2025-10-01 | 2814 |
| 2025-11-01 | 2360 |
| 2025-12-01 | 1816 |
| 2026-01-01 | 2107 |
| 2026-02-01 | 2970 |
| 2026-03-01 | 3817 |
| 2026-04-01 | 3259 |
| 2026-05-01 | 3282 |
| 2026-06-01 | 3023 |
| 2026-07-01 | 3512 |
Data table
| Date | Input processed, cumulative (EB) |
|---|---|
| 2023-10-31 | 0.005 |
| 2023-11-30 | 0.007 |
| 2023-12-31 | 0.01 |
| 2024-01-31 | 0.013 |
| 2024-02-29 | 0.019 |
| 2024-03-31 | 0.025 |
| 2024-04-30 | 0.033 |
| 2024-05-31 | 0.061 |
| 2024-06-30 | 0.112 |
| 2024-07-31 | 0.198 |
| 2024-08-31 | 0.256 |
| 2024-09-30 | 0.341 |
| 2024-10-31 | 0.408 |
| 2024-11-30 | 0.489 |
| 2024-12-31 | 0.522 |
| 2025-01-31 | 0.579 |
| 2025-02-28 | 0.644 |
| 2025-03-31 | 0.678 |
| 2025-04-30 | 0.699 |
| 2025-05-31 | 0.73 |
| 2025-06-30 | 0.763 |
| 2025-07-31 | 0.798 |
| 2025-08-31 | 0.833 |
| 2025-09-30 | 0.867 |
| 2025-10-31 | 0.92 |
| 2025-11-30 | 0.983 |
| 2025-12-31 | 1.054 |
| 2026-01-31 | 1.132 |
| 2026-02-28 | 1.205 |
| 2026-03-31 | 1.302 |
| 2026-04-30 | 1.407 |
| 2026-05-31 | 1.506 |
| 2026-06-30 | 1.596 |
| 2026-07-31 | 1.708 |
| 2026-08-31 | 1.797 |
DiracX / CWL integration
Python · CWL · DiracX
DIRAC has scheduled LHCb's grid workloads for two decades, and its job descriptions have always been bespoke, closed to the wider ecosystem of workflow tooling. For DiracX, its successor, we are adopting the Common Workflow Language, so that a physics workload is described in an open, portable standard rather than an in-house dialect.
I presented our strategy at CHEP 2026, as Aligning DIRAC Workflows with CWL.
Autonomous operations service
Python · FastAPI · Celery · MySQL
Operating a system that runs LHCb's offline processing means diagnosing job failures at a rate no person can sustain by reading logs. What existed was a log-analysis script that worked well for exactly one careful expert at a time: genuinely useful, but unshareable, and with nothing to stop a mistyped action doing real damage.
I turned it into a concurrency-safe service a team of operators can work with: adapting the root-cause grouping across thousands of jobs, and exposing it clearly through an operations dashboard, and a carefully commissioned-action lifecycle in which proposed interventions require operator approval, run through guarded executors, and are continuously reconciled against the actual state of the system. Nothing acts unsupervised, and everything leaves an audit trail.
LbAgents: a proof-of-concept agentic toolkit for LHCb Computing Operations
Python · LangGraph · MCP · OpenSearch · Langfuse
Operations generates a lot of repetitive diagnostic work: a CI job fails, somebody reads the log, recognises the shape of the problem, and writes much the same advice they wrote last month. LbAgents is a set of LangGraph workflows that take the first pass: diagnosing CI failures, with single-agent and multi-agent variants, investigating stuck productions, and answering support questions arriving over Mattermost.
The design decisions matter more here than the choice of model or even the workflows themselves (they will require optimisations regardless). Graph factories return uncompiled graphs, so the host application owns checkpointing. Tools arrive as parameters rather than hardcoded, discovered over MCP at runtime, so adding a workflow needs no change in the host. Every write tool, whether that is posting a review or opening an issue, sits behind a human-in-the-loop interrupt that shows the pending call to a person before it runs. Nothing an agent decides reaches GitLab unreviewed.
lbfluka
Python · LHCbDIRAC · FLUKA · pixi
FLUKA is a long-established Fortran Monte Carlo code, and running a campaign with it used to mean hand-maintained shell scripts: build the binary, stage the data, launch the jobs, collect whatever comes back. That works for one person running one study. It does not survive thousands of grid jobs, several people, and a result somebody will want to reproduce in two years.
lbfluka allows one to express this workflow instead as DIRAC transformations while carefully adapting to FLUKA's constraints. A campaign is declared in a YAML file and submitted from a merge request; the tooling packages the compiled binary and its data, fans the work out across thousands of grid nodes with deterministic per-job seeds, merges the outputs, and registers the results in the bookkeeping so every file can be traced back to the configuration that produced it. The physics is FLUKA's. The engineering is making it reproducible at scale, and the same shape as every other production the collaboration runs.
LbAPWeb
React · TypeScript
The web application at the centre of the Run 3 data-processing workflow: where analysts watch their CI tests, debug failing jobs, browse their n-tuples interactively, and manage samples through their lifecycle. I built it during my PhD to replace the previous interface, motivated directly by my own experience of validating productions as an analyst.
It has since been adopted well beyond its original scope, including as the frontend for LHCb's simulation-request system. Resource-usage plots introduced through this work have led to multiple memory-leak discoveries in production workflows, and built-in n-tuple browsing means checking an output no longer requires downloading it.