Week 29 · September 2026

44 Person-Years of Proofreading: The Complete Male Fly Connectome

September 13, 2026 · by Satish K C 10 min read
Connectomics Computer Vision Segmentation Datasets

The Paper

"Sexual dimorphism in the complete connectome of the Drosophila male central nervous system" was published in Cell on September 3, 2026. Stuart Berg and Isabella R. Beckett are the first authors. Gregory S. X. E. Jefferis of the MRC Laboratory of Molecular Biology in Cambridge, Gerald M. Rubin of Janelia, and Berg are the corresponding authors. The work comes from the FlyEM team at HHMI Janelia, the Cambridge Connectomics Group and Google Research. Google Research published its own account of the reconstruction stack the same week, written by Michal Januszewski and Viren Jain. The claim is a complete wiring diagram of one male fruit fly central nervous system. That is 166,700 neurons and 11,710 cell types. It covers the central brain, the optic lobes and the ventral nerve cord in a single imaged volume, with the neck connective intact. By neuron count it is the largest complete brain map published so far. It also delivers the first male-to-female comparison of a nervous system at synaptic resolution.

The Problem Before This Paper

Connectomics had pieces, not a whole animal. The 2020 hemibrain covered roughly 25,000 neurons and 21 million connections, which is about half of one central brain. The 2024 FlyWire release covered a complete adult female brain at 139,255 neurons and 54.5 million synapses. The male ventral nerve cord had its own separate dataset, MANC. None of these came from the same animal. None contained a male brain. So every claim about sex differences in fly wiring was a comparison across datasets that differed in imaging modality, voxel size, segmentation model and proofreading policy. FlyWire was imaged by serial-section transmission electron microscopy at 4 x 4 x 40 nm. The hemibrain was focused ion beam SEM at 8 nm isotropic. When two such datasets disagree about a neuron, you cannot tell whether the animal differs or the pipeline does. The second gap was the neck. Brain and nerve cord datasets were physically cut apart. That severs the descending and ascending pathways at exactly the point where a decision turns into movement.

What They Built

Seven enhanced focused ion beam scanning electron microscopy systems ran for 13 months. They produced one volume at 8 x 8 x 8 nm isotropic resolution: 160 teravoxels covering 0.082 cubic millimetres of tissue. Neurons were segmented automatically with flood-filling networks, the recurrent convolutional method Januszewski and colleagues introduced in 2018. A flood-filling network starts inside one neurite and extends outward along it, rather than labelling voxels independently. Synapses were detected automatically at average precision and recall of 0.82 and 0.81. That produced 46 million presynapses connected to 312 million postsynaptic densities. Humans then spent an estimated 44 person-years correcting the result. Proofreading was prioritised, not uniform. Every fragment with more than 100 synaptic connections was reviewed. Coverage was audited against independently detected cell nuclei, and 98.9% of 141,780 nuclei matched a proofread neuron. Cell typing combined NBLAST morphology scores with connectivity co-clustering, arranged as a four-level hierarchy of superclass, hemilineage, supertype and type. 97.5% of neurons were matched to a counterpart in FlyWire, the hemibrain or MANC. Sex-determination gene expression was assigned by co-registering light microscopy datasets, giving 4,505 fruitless neurons and 407 doublesex neurons, with 251 cells expressing both.

Flood-filling network Starts inside one neurite and extends outward along it, rather than labelling voxels independently. seed 0.5 um field of view Small field of view keeps merge errors low. The cost is splits, which is what the next two stages exist to repair.

The segmentation grows along one process at a time. Neighbouring neurites in the same field of view are what a merge error looks like, and a single merge corrupts every path through the graph.

Male CNS reconstruction pipeline Four automated stages, one manual stage. The manual stage is the one that does not scale. 01 IMAGING 7 x eFIB-SEM 13 months 8 nm isotropic 160 teravoxels 02 SEGMENT Flood-filling networks recurrent CNN, grows along neurite 03 SYNAPSES 46M presynapses 312M PSDs precision 0.82 recall 0.81 04 PROOFREAD 44 person-years by hand fragments above 100 connections 05 TYPE 166,700 neurons 11,710 cell types Fully automated - cost scales with compute Manual - scales with people PUBLISHED COMPLETION RATES 94% presynaptic 42% postsynaptic 40.1% of connections have both partners proofread

Segmentation and synapse detection are fully automated. Verification is not. The completion rates are reported per stage rather than folded into one accuracy number.

The headline synapse count needs one line of arithmetic before you use it. Google Research reports 125 million synaptic connections. The paper reports 312 million postsynaptic densities. Both are correct, and the gap between them is a quality filter:

312M detected PSDs × 40.1% with both partners proofread ≈ 125M usable connections

312 million is the raw detector output. 125 million is the subset where both the sending and the receiving neuron passed human review. The first number describes the model. The second number describes what you can safely build an analysis on. Publishing both, and the 94% versus 42% pre-to-post asymmetry alongside them, is the part of this release worth copying.

Key Findings

Results

The scale numbers are best read against the datasets they replace. Per-neuron proofreading cost is the line that matters, because it is the only one that tells you whether the method is improving or the budget is:

Dataset Year Coverage Neurons Synapses Imaging Proofreading
Hemibrain2020Half central brain, female~25,000 ~21MFIB-SEM 8 nm iso~50 person-years
FlyWire / FAFB2024Whole brain, female139,255 54.5MssTEM 4 x 4 x 40 nm33 person-years
MICrONS20251 mm³ mouse visual cortex>200,000 cells ~0.5BssEM, ~4 nmPartial
Male CNS2026Whole CNS, male166,700 125M usable / 312M PSDseFIB-SEM 8 nm iso44 person-years

MICrONS is larger in synapse count but is a one cubic millimetre sample of cortex, not a complete nervous system. The male CNS is the largest complete one.

Per neuron, the hemibrain cost about 4 hours of human review. The male CNS cost about 32 minutes. That is a 7.5 times improvement in unit cost across six years of better models. Total cost still went up in absolute terms, from 50 person-years to 44 for 6.7 times the neurons, only because the automation improvement barely outran the volume increase. On the biology, sex differences turn out to be small, localised and high in the hierarchy. 8,069 cell types are shared. 498 are not. Roughly 5% of the male central brain carries the difference and the other 95% is common wiring. Named examples are concrete: AOTU008 carries ventral axonal projections that are absent in the female homologue and commits 57% of its output to sex-specific or dimorphic partners. LoVP92 is the only male-specific visual projection neuron, about six cells per hemisphere, with 18% of inputs and 43% of outputs on dimorphic partners. vpoEN is morphologically identical in both sexes and rewires downstream, sending signal to vpoDN in females and P1_6a in males.

Where the sexes differ 8,567 matched cell types. The behavioural difference is total. The wiring difference is not. 8,069 isomorphic 94.2% of matched types are shared between male and female THE 498 THAT DIFFER 138 dimorphic 289 male-specific 71 female-specific SHARE OF THE CENTRAL BRAIN 4.8% male 2.4% female concentrated in higher brain centres NAMED IN THE PAPER AOTU008 57% of output to sex-specific LoVP92 18% in, 43% out, dimorphic vpoEN same shape, different targets

The small blocks are drawn to scale in the top bar and then magnified, because at true scale the 498 types that differ are almost invisible against the 8,069 that do not.

Why This Matters for AI and Automation

Three-stage reconstruction: a 0.5 micrometre model splits a neurite into fragments, a 4 micrometre model proposes ways to join them, and a 20 micrometre model rejects the implausible assemblies and keeps one neuron. Reported at 94.2% normalised expected run length against about 52% for the prior method, and 67,200 cubic micrometres per hour against a prior estimate of 800.

PATHFINDER in three stages. The small model is deliberately conservative and leaves splits behind. The two larger-context models exist to repair them, which is the same division of labour as a candidate generator followed by a critic.

The scaling arithmetic

Why the mouse brain is a method problem, not a funding problem

This connectome covers 0.082 mm³. A mouse brain is roughly 500 mm³, about 6,100 times larger. At the same proofreading rate that is on the order of 270,000 person-years. No budget fixes that.

Apply PATHFINDER's claimed 84 times throughput improvement and it drops to roughly 3,200 person-years. Better, and still about 70 times the entire effort behind this paper. Two more order-of-magnitude improvements are needed before a mouse brain is a project rather than a thought experiment, and the human brain is another 516,000 times the neuron count from here.

My Take

The headline says AI mapped a brain. The numbers say something more useful. Automated segmentation handled 160 teravoxels without human help, and the project still needed 44 person-years of people looking at neurons. That is not a failure of the models. It is what progress actually looks like when a system is judged on completeness rather than on an average. A 0.82 precision synapse detector is excellent and also useless on its own, because a connectome is a graph where one bad merge corrupts every path through it, so the tolerable error rate is set by the longest path you care about and not by the mean. The honest read on PATHFINDER is that it is the right response and it is not yet in this result. Its 84 times figure comes from exhaustive axon reconstruction in mouse cortex, not from a shipped fly connectome, and the paper's own limitations are the ones that matter: ground truth is still assembled by hand, SHAPE accuracy scales only logarithmically with training data, and the inference-time trick that cuts F1 error by up to 44% costs 16 rotations times 16 point samples per candidate pair. That is a real technique with a real bill attached. What I find most interesting is the biology as a controlled experiment. Two animals, same species, same body plan, one variable. 8,069 shared cell types and 498 that are not. The difference in behaviour is total and the difference in wiring is 5%, concentrated in higher centres, with the sensory and motor periphery untouched. Anyone who has watched a small targeted change to a model produce a large change in behaviour will recognise the shape. The difference is that here somebody counted it exactly, then published the 42% number that makes their own dataset look weaker.

Discussion question: This project automated the expensive stage completely and the verification stage not at all, and the verification stage is now the entire cost. If you audited your own AI pipeline the way these authors audited theirs, reporting completion separately for every stage instead of one end-to-end score, which stage would turn out to be carrying the real cost, and would that number change what you automate next?

← Back to all papers
Share