New Nine application domains — see what we are tracking
Understanding the systems we are building.
We study how machine learning systems represent the world, where those representations break down, and what it takes to make model behaviour predictable enough to rely on.
- Research areas
- 4
- Application domains
- 9
Interpretability, measurement, robustness, efficiency
From robotics and genomics to formal mathematics
Our current work
Four questions we keep
returning to
Our agenda is narrow on purpose. Each area exists because we think the field is measuring something other than what it claims to measure.
01
Interpretability
Neural networks build internal representations nobody designed. We work on methods that read those representations directly, so a claim about why a model did something can be tested rather than asserted.
Read more02
Evaluation & measurement
A benchmark score is a measurement, and measurements drift. We study how evaluations decay as models are trained against them, and build tests whose results still mean something after the number has been optimised.
Read more03
Robustness & alignment
Systems that behave well in evaluation often behave differently in deployment. We look at behaviour under distribution shift, adversarial pressure and open-ended use, and at which training signals actually generalise.
Read more04
Efficient architectures
Capability per unit of compute is a research question, not just an engineering one. We investigate sparsity, retrieval and long-context memory, with an interest in designs that are easier to inspect as well as cheaper to run.
Read moreMethod
Read the representation,
not the output
A model’s answer tells you what it did. Its internal state tells you why. We decompose activations into a sparse basis and ask whether the features we recover are the ones the model is actually using — or the ones our method happens to be good at finding.
How we work
Fewer results, held to a higher bar
Publish the negative results
Most of what a lab learns is which approaches do not work. We write those up too. A research programme that only reports its successes is not reporting its findings.
Release the artefacts
Code, evaluation suites and model probes ship with the paper under a permissive licence. A result that cannot be reproduced by a reader is a claim, not a finding.
Say what we do not know
Every write-up carries an explicit limitations section stating the conditions under which we expect the result to fail. We would rather be narrow and correct than broad and unfalsifiable.
Open tooling
Every figure here is
a reproducible artefact
We release the harness that produced a result alongside the result itself. If a reader cannot regenerate the plot on their own machine, we have published a claim rather than a finding — and we would rather not do that.
Where we intend to apply it
Nine domains we are
looking to work in
These are not active projects. They are the fields whose measurement problems we think are the most underserved — and where we are looking for collaborators who already have the data and the problem.
Robotics & embodied learning
Vision-language-action policies, sim-to-real transfer, dexterous manipulation.
Genomics & protein design
Structure prediction beyond single chains, generative design, variant effects.
Mathematics & formal reasoning
Neural theorem proving, autoformalisation, machine-checked correctness.
Materials & chemistry
Interatomic potentials, crystal stability, retrosynthesis, autonomous labs.
Climate & earth systems
Learned forecast emulators, extremes, downscaling, emissions monitoring.
Neuroscience & interfaces
Speech decoding, representational alignment, decoder drift, connectomics.
Clinical decision support
Multimodal clinical models, shift between hospitals, prospective evaluation.
Software & program synthesis
Repository-scale agents, verification, long-horizon reliability.
Multi-agent & mechanism design
Learning agents in markets, robust mechanisms, policy simulation.
Bring us a
measurement problem
If you work in one of these domains and cannot get traction on knowing whether your system actually works, that is the collaboration we are looking for.