Projects

LHCb Analysis Productions

Python · FastAPI · React · TypeScript · Celery · MySQL · SQLAlchemy · Kubernetes

Data collected by LHCb requires further processing to be useful for physics analysis, which depends on the kind of physics an analyst is looking for and/or trying to measure.

Four stages left to right: detector output, huge and needing filtering; an analyst asking 'where's my decay?'; the code they wrote, which finds, filters and calculates what they need; and the resulting dataset, a table with one row per decay candidate.
The LHCb analyst user story: what every LHCb analysis needs before the physics can start.

Not too long ago, an LHCb analyst would produce an analysis dataset by submitting their own jobs to the Worldwide LHC Computing Grid (WLCG) or simply the Grid (a collection of computing centres which execute jobs a bit like in a batch system), and effectively babysitting them: monitoring progress, diagnosing failures job by job, resubmitting, and preserving the outputs and configuration by hand. It was the most time-consuming and error-prone part of starting an analysis, and when LHC Run 3 multiplied the data rates this became completely unworkable.

Four stages left to right: an analyst submitting their own jobs; those jobs running at Grid computing sites, where two of the nine have failed because untested configurations waste the allocated resources; the analyst babysitting them, watching progress and resubmitting while failures often go ignored; and the outputs and configuration kept by hand, as a stack of papers labelled 'v3_final?'.
Before Analysis Productions: submit, watch, resubmit, and keep track of it all yourself.

Analysis Productions replaces that with a declarative, centrally managed service: the analyst describes what they need in a short YAML file, and the platform handles validation, CI testing (using real Grid jobs), submission, monitoring, and preservation, with provenance kept end to end so that any dataset can be traced back to exactly how it was made.

A short YAML file, in which the analyst describes what they need rather than how to run it, feeds a centrally managed service that validates, tests in CI with real Grid jobs, submits, monitors and preserves; out comes the dataset, produced and preserved, ready to analyse.
The analyst declares what they need; the service and operators handle the rest.

The platform grew out of LHCb's earlier working-group productions system and was already running when I joined the collaboration back in 2020. Since then, I have worked on much of the API backend and (completely new, overhauled) web frontend LbAPWeb (see below), and I remain one of the platform's core developers and maintainers.

It has been hugely successful. Analysis Productions is the official processing route for LHCb's Run 3 data, used by hundreds of analysts and open to the whole collaboration. As of August 2026 it had read 1.8 exabytes of collision data and returned 7.6 PB of analysis-ready datasets, a reduction of more than 200 to 1. Chris Burr presented the platform at CHEP 2026, as an exascale analysis data processing and management service for LHCb.

Analysis Productions samples created per month
05001,0001,5002,0002,5003,0003,5004,000↑ Samples created / month20202021202220232024202520262020-01-01: 422020-02-01: 1332020-03-01: 1882020-04-01: 3142020-05-01: 1002020-06-01: 1172020-07-01: 1342020-08-01: 282020-09-01: 462020-10-01: 662020-11-01: 882020-12-01: 1982021-01-01: 2862021-02-01: 2712021-03-01: 2582021-04-01: 3902021-05-01: 5462021-06-01: 702021-07-01: 1,1752021-08-01: 1262021-09-01: 362021-10-01: 3142021-11-01: 1922021-12-01: 962022-01-01: 1442022-02-01: 2452022-03-01: 5162022-04-01: 2062022-05-01: 6062022-06-01: 5252022-07-01: 802022-08-01: 2262022-09-01: 5922022-10-01: 3722022-11-01: 8362022-12-01: 1582023-01-01: 2412023-02-01: 4442023-03-01: 1,4002023-04-01: 8312023-05-01: 3632023-06-01: 5112023-07-01: 3932023-08-01: 3682023-09-01: 1,2362023-10-01: 1,6822023-11-01: 3812023-12-01: 6972024-01-01: 7202024-02-01: 8882024-03-01: 7922024-04-01: 1,3592024-05-01: 1,1412024-06-01: 9572024-07-01: 1,5932024-08-01: 1,0372024-09-01: 1,3082024-10-01: 1,4142024-11-01: 2,4322024-12-01: 1,2382025-01-01: 1,5612025-02-01: 3,5092025-03-01: 1,3992025-04-01: 9392025-05-01: 1,7762025-06-01: 1,5352025-07-01: 2,8062025-08-01: 1,7052025-09-01: 1,7792025-10-01: 2,8142025-11-01: 2,3602025-12-01: 1,8162026-01-01: 2,1072026-02-01: 2,9702026-03-01: 3,8172026-04-01: 3,2592026-05-01: 3,2822026-06-01: 3,0232026-07-01: 3,512
One sample per transformation family, counted in the month of its first transformation.
Data table
DateSamples created / month
2020-01-0142
2020-02-01133
2020-03-01188
2020-04-01314
2020-05-01100
2020-06-01117
2020-07-01134
2020-08-0128
2020-09-0146
2020-10-0166
2020-11-0188
2020-12-01198
2021-01-01286
2021-02-01271
2021-03-01258
2021-04-01390
2021-05-01546
2021-06-0170
2021-07-011175
2021-08-01126
2021-09-0136
2021-10-01314
2021-11-01192
2021-12-0196
2022-01-01144
2022-02-01245
2022-03-01516
2022-04-01206
2022-05-01606
2022-06-01525
2022-07-0180
2022-08-01226
2022-09-01592
2022-10-01372
2022-11-01836
2022-12-01158
2023-01-01241
2023-02-01444
2023-03-011400
2023-04-01831
2023-05-01363
2023-06-01511
2023-07-01393
2023-08-01368
2023-09-011236
2023-10-011682
2023-11-01381
2023-12-01697
2024-01-01720
2024-02-01888
2024-03-01792
2024-04-011359
2024-05-011141
2024-06-01957
2024-07-011593
2024-08-011037
2024-09-011308
2024-10-011414
2024-11-012432
2024-12-011238
2025-01-011561
2025-02-013509
2025-03-011399
2025-04-01939
2025-05-011776
2025-06-011535
2025-07-012806
2025-08-011705
2025-09-011779
2025-10-012814
2025-11-012360
2025-12-011816
2026-01-012107
2026-02-012970
2026-03-013817
2026-04-013259
2026-05-013282
2026-06-013023
2026-07-013512
Data processed by Analysis Productions
0.00.20.40.60.81.01.21.41.61.8↑ Input processed, cumulative (EB)Jan2024AprJulOctJan2025AprJulOctJan2026AprJul2023-10-31: 0.005 EB2023-11-30: 0.007 EB2023-12-31: 0.01 EB2024-01-31: 0.013 EB2024-02-29: 0.019 EB2024-03-31: 0.025 EB2024-04-30: 0.033 EB2024-05-31: 0.061 EB2024-06-30: 0.112 EB2024-07-31: 0.198 EB2024-08-31: 0.256 EB2024-09-30: 0.341 EB2024-10-31: 0.408 EB2024-11-30: 0.489 EB2024-12-31: 0.522 EB2025-01-31: 0.579 EB2025-02-28: 0.644 EB2025-03-31: 0.678 EB2025-04-30: 0.699 EB2025-05-31: 0.73 EB2025-06-30: 0.763 EB2025-07-31: 0.798 EB2025-08-31: 0.833 EB2025-09-30: 0.867 EB2025-10-31: 0.92 EB2025-11-30: 0.983 EB2025-12-31: 1.054 EB2026-01-31: 1.132 EB2026-02-28: 1.205 EB2026-03-31: 1.302 EB2026-04-30: 1.407 EB2026-05-31: 1.506 EB2026-06-30: 1.596 EB2026-07-31: 1.708 EB2026-08-31: 1.797 EB
Cumulative input read by Analysis Productions jobs. Counts successful jobs only, and a file read by several jobs counts each time.
Data table
DateInput processed, cumulative (EB)
2023-10-310.005
2023-11-300.007
2023-12-310.01
2024-01-310.013
2024-02-290.019
2024-03-310.025
2024-04-300.033
2024-05-310.061
2024-06-300.112
2024-07-310.198
2024-08-310.256
2024-09-300.341
2024-10-310.408
2024-11-300.489
2024-12-310.522
2025-01-310.579
2025-02-280.644
2025-03-310.678
2025-04-300.699
2025-05-310.73
2025-06-300.763
2025-07-310.798
2025-08-310.833
2025-09-300.867
2025-10-310.92
2025-11-300.983
2025-12-311.054
2026-01-311.132
2026-02-281.205
2026-03-311.302
2026-04-301.407
2026-05-311.506
2026-06-301.596
2026-07-311.708
2026-08-311.797

DiracX / CWL integration

Python · CWL · DiracX

DIRAC has scheduled LHCb's grid workloads for two decades, and its job descriptions have always been bespoke, closed to the wider ecosystem of workflow tooling. For DiracX, its successor, we are adopting the Common Workflow Language, so that a physics workload is described in an open, portable standard rather than an in-house dialect.

I presented our strategy at CHEP 2026, as Aligning DIRAC Workflows with CWL.

Autonomous operations service

Python · FastAPI · Celery · MySQL

Operating a system that runs LHCb's offline processing means diagnosing job failures at a rate no person can sustain by reading logs. What existed was a log-analysis script that worked well for exactly one careful expert at a time: genuinely useful, but unshareable, and with nothing to stop a mistyped action doing real damage.

I turned it into a concurrency-safe service a team of operators can work with: adapting the root-cause grouping across thousands of jobs, and exposing it clearly through an operations dashboard, and a carefully commissioned-action lifecycle in which proposed interventions require operator approval, run through guarded executors, and are continuously reconciled against the actual state of the system. Nothing acts unsupervised, and everything leaves an audit trail.

LbAgents: a proof-of-concept agentic toolkit for LHCb Computing Operations

Python · LangGraph · MCP · OpenSearch · Langfuse

Operations generates a lot of repetitive diagnostic work: a CI job fails, somebody reads the log, recognises the shape of the problem, and writes much the same advice they wrote last month. LbAgents is a set of LangGraph workflows that take the first pass: diagnosing CI failures, with single-agent and multi-agent variants, investigating stuck productions, and answering support questions arriving over Mattermost.

The design decisions matter more here than the choice of model or even the workflows themselves (they will require optimisations regardless). Graph factories return uncompiled graphs, so the host application owns checkpointing. Tools arrive as parameters rather than hardcoded, discovered over MCP at runtime, so adding a workflow needs no change in the host. Every write tool, whether that is posting a review or opening an issue, sits behind a human-in-the-loop interrupt that shows the pending call to a person before it runs. Nothing an agent decides reaches GitLab unreviewed.

lbfluka

Python · LHCbDIRAC · FLUKA · pixi

FLUKA is a long-established Fortran Monte Carlo code, and running a campaign with it used to mean hand-maintained shell scripts: build the binary, stage the data, launch the jobs, collect whatever comes back. That works for one person running one study. It does not survive thousands of grid jobs, several people, and a result somebody will want to reproduce in two years.

lbfluka allows one to express this workflow instead as DIRAC transformations while carefully adapting to FLUKA's constraints. A campaign is declared in a YAML file and submitted from a merge request; the tooling packages the compiled binary and its data, fans the work out across thousands of grid nodes with deterministic per-job seeds, merges the outputs, and registers the results in the bookkeeping so every file can be traced back to the configuration that produced it. The physics is FLUKA's. The engineering is making it reproducible at scale, and the same shape as every other production the collaboration runs.

LbAPWeb

React · TypeScript

The web application at the centre of the Run 3 data-processing workflow: where analysts watch their CI tests, debug failing jobs, browse their n-tuples interactively, and manage samples through their lifecycle. I built it during my PhD to replace the previous interface, motivated directly by my own experience of validating productions as an analyst.

It has since been adopted well beyond its original scope, including as the frontend for LHCb's simulation-request system. Resource-usage plots introduced through this work have led to multiple memory-leak discoveries in production workflows, and built-in n-tuple browsing means checking an output no longer requires downloading it.