The Attending Pattern: A Frontier LLM Orchestrating Specialized Tools on Patient-Owned Longitudinal CT
Daniel Uribe, founder and CEO, GenoBank.io
GenoBank Research Team
June 26, 2026
A build report. What was constructed, what it cost in time, tokens, and GPU, and where each layer reached its honest limit.
Processing identity (operator biowallet, EIP-55): 0x5f5a60Ea…d19a. The data subject in the case study is referenced pseudonymously as the index case; it is not the author.
1. Motivation
A cancer patient on a clinical trial accumulates serial CT studies over months. The clinically important question is rarely "what is in this scan" in isolation. It is "what changed since last time, and does it matter." Answering it well requires registration across timepoints, change detection, measurement under a response standard, and, above all, judgment about which numbers are trustworthy. Automated imaging tools are improving fast at the first three. They remain weak, and frequently overconfident, at the fourth.
The motivating event for this report is concrete. An internal change-detection pipeline, run on a patient-owned vault, flagged a new lung lesion of roughly 12 mm and issued a RECIST call of progressive disease on a three-month chest interval. A reading radiologist on the same study described instead a new 3 mm ill-defined nodule, favored as a small focus of infection, with no progressively growing suspicious nodule. One of these reads, if acted on, would change a treatment plan. They cannot both be right. Resolving the discrepancy honestly, and understanding why the automated tool was confident and wrong, is the subject of this paper.
The system is built on a simple ownership premise that GenoBank.io has carried across genomics and now extends to imaging: the patient holds the asset. Each imaging study is acquired into a vault the patient controls, registered as a content-addressed biological object (a BioCID) with on-chain anchoring and revocable consent, and never copied to an analyst's laptop. The heavy bytes stay server-side; only small derived results travel.
2. Architecture
Every operation, acquisition, registration, segmentation, change detection, response measurement, and erasure, is a verb of the biofs protocol ("biological filesystem"). A verb is a thin client that submits a job manifest to the biofs-node orchestrator. The orchestrator is the only component that mounts cloud storage, schedules GPU, anchors provenance on-chain, and assembles a result manifest. This keeps the system reproducible: a result is reproducible by typing a verb, not by re-running a one-off script.
The LLM does not replace this layer. It calls it. The LLM is the orchestrator and reader of the orchestrator's outputs, the layer we describe as the attending.
2.1 The biofs imaging verbs
| Verb | Function | Compute |
|---|---|---|
imaging acquire (eUnity sync) | Pull a DICOM study from a hospital viewer into the patient vault, server-side, deduplicated by Study Instance UID, anchored as a BioCID. | Server-side I/O |
imaging compare | Image Time Machine. Rigid, deformation-free slice-by-slice comparison of two timepoints, with blink, crossfade, and signed-difference views and an optional z-band region crop. Rigid only, so a real focal change is never warped away. | CPU |
imaging twin | N-timepoint registered organ meshes (a strict-isometry digital twin) for volume trends across studies. | GPU |
imaging findings | Tumor detection and longitudinal lesion tracking with a response read, bounded to the detector's classes. | GPU |
imaging characterize | Per-candidate feature panel (engineered radiomics and several learned backbones) plus a vision-language read and a lesion-versus-artifact head. | GPU |
imaging lesions (volumetric RECIST) | Independent per-timepoint segmentation, cross-timepoint linking, true RECIST category, and volume doubling time, to catch iso-attenuating growth a density subtraction misses. | GPU |
imaging trajectory | Per-voxel persistent-versus-transient classifier across three timepoints, to separate a monotonic change from a one-interval blip. | CPU |
imaging attribute | Frame-exact organ labeling of each detected focus by whole-body segmentation, so every change carries an organ tag and a confidence. | GPU |
imaging enrich | Foundation-model overlays (a CT embedding model, a medical vision-language model, and a tumor segmenter) attached to a comparison. | GPU |
2.2 The imaging-radiology MCP server
A Model Context Protocol server exposes the orchestrator's outputs to the LLM as read tools, so the model reasons over structured facts rather than free narration. It carries the standards it applies (RECIST 1.1 measurability rules, response categories, and the ACR and RSNA structured-report sections) with citations, so any read is transparent and grounded. Tools include the comparison listing and the raw change analysis, a deterministic RECIST read and a per-lesion volumetric RECIST read, an organ-attribution read, a longitudinal pathogenic-findings read, an ACR and RSNA structured report assembler, the applied-standards reference, and four literature tools for surveying the field. Every output is framed as decision support, not a diagnosis, and confidence is graded and reported rather than withheld.
2.3 The clinical-records spine (Epic FHIR R4)
Imaging without the clinical record is half a picture. The same vault holds the patient's Epic records, pulled through a SMART-on-FHIR integration the patient authorizes. During this work we repaired four real defects in that integration and verified the repairs end to end:
- An authorization redirect that appeared to loop. The token exchange and a multi-minute synchronous import ran inside one callback, which exceeded the content delivery network's proxy timeout, so the browser callback timed out and fell back to the hospital login while the import quietly completed. The fix moves the import to a background thread and redirects immediately; the dashboard then polls for status. The repaired flow returns in roughly two seconds instead of nine to twelve minutes.
- Report text that was never captured. The import stored pointers to report and note bodies, not the text. Radiology impressions arrive as Binary references on the diagnostic report, not inline. We added a best-effort pass that reads the just-imported diagnostic-report and document bundles back from cache, resolves the Binary references while the token is live, decodes the HTML, and stores the narrative. This is what made the radiologist's own words readable, the ground truth this paper depends on.
- A silent truncation cap. Two large observation categories were being cut at a per-resource time budget. With a token budget of roughly fifty-five minutes against an import that uses about twelve, the caps were leaving data on the table for no reason. Raising them let those categories page out fully. A representative full import returned 23,949 records across 36 of 51 resource types, with all radiology narratives dereferenced.
The fourth defect was a dashboard button that started the authorization flow. It silently returned when no hospital was selected from a dropdown that defaulted to an empty placeholder, so a click produced nothing and reached no server endpoint. The fix auto-selects the patient's primary hospital and falls back to the first real option, so the flow is a single click.
3. Case study: a contested chest interval
We applied the pipeline to an index oncology case, a patient with pancreatic adenocarcinoma after a Whipple resection, on a clinical trial, with serial chest and abdomen CT. The data subject is consented and pseudonymous. The contest is the one described in the motivation: an automated read of new 12 mm lesion and progressive disease, against a radiologist read of a new 3 mm indeterminate nodule and no progression.
3.1 Reconciliation
An orchestrated workflow of five independent reasoning agents, each grounded in the pipeline's actual job outputs, converged on a single verdict: the automated 12 mm finding and the radiologist's 3 mm nodule are different objects, not a mis-sizing of the same object. The evidence was convergent, with several axes independently disqualifying.
| Axis | Radiologist 3 mm nodule | Pipeline flagged foci |
|---|---|---|
| Location | Posterior right upper lobe, high thorax | Lung base near the diaphragm, opposite end of the field of view |
| Organ attribution | Lung parenchyma | Unsegmented or bone, stomach, aorta, and colon at the field-of-view edge, confidence near zero, none in lung |
| Size | About 3 mm, roughly 0.014 mL | Tens of mL, a thousandfold mismatch |
| Mechanism | Discrete focus | Large appeared blobs paired with co-located resolved blobs at the same slices, the signature of tissue shifting between scans |
| Engine self-grade | Not applicable | Progression signal at low confidence, high regional motion burden near 986 mL, lowest bone-overlap quality of any chest pair |
The 12 mm in the automated headline was the long axis its lesion-tracking path reported for the new lesion. The separate change-detection job surfaced the same lung-base motion as much larger appeared blobs, on the order of 30 to 50 mm long axis. Both are the one artifact seen at two processing stages. The true 3 mm nodule, meanwhile, was never surfaced by the change job at all. At roughly 0.014 mL it sits more than twentyfold below that job's 0.3 mL reporting floor, and under its downsampled detection grid. Its absence is expected at that size and is not evidence against the radiologist. The automated progression call was a registration and respiratory-motion artifact at the deformable lung base, and should not have reached a clinician as progression.
3.2 Looking at the pixels directly
We then did the thing that, on its face, sounds like it should make the specialized tools unnecessary: the model rendered the exact slices from the vault DICOM and looked at them. Series 4, the 0.6 mm lung-kernel reconstruction, at the cited image and its neighbors. Several window settings. The March study beside the June study. Finally a native-resolution registered comparison: the model rigidly aligned the full March volume into the June grid (correcting an absolute table-position offset of roughly 200 mm) and rendered a same-voxel difference at the nodule level.
The honest result. The model could not confirm the 3 mm nodule. In the registered difference the bright structures were vessel-edge residual from breathing, which rigid registration cannot undo, and which at the 3 mm scale both masks and mimics a real focus. What the model could do was exclude any gross or large new lesion, which independently corroborated that the 12 mm progression call was not a real mass, and note that a mediastinal window showed no solid soft-tissue dot, weakly consistent with the radiologist's lean toward an infectious, non-solid focus. The model was bound by the same physical limit as the tools. A 3 mm nodule is sub-measurable under the response standard, near the partial-volume floor of the scanner, and below the reliable detection threshold of every automated detector in the stack. Respecting that limit honestly is different from beating it. The radiologist found the nodule. The model did not.
The abdomen and pelvis study from the same date, read by the institution's radiologist, showed no evidence of recurrent disease after the Whipple resection, no lymphadenopathy, and no ascites. The combined restaging picture was therefore reassuring, with the one indeterminate finding a tiny nodule favored as infectious, to be settled by a short-interval follow-up CT.
4. Cost
We report cost honestly, because the most useful single fact here is an efficiency one. The entire reconciliation and direct-imaging analysis used no GPU. The comparison and the native registration are CPU jobs on commodity server hardware. The GPU-dependent perception verbs (the digital twin, detection, characterization, and volumetric RECIST) run on an A100 with an idle auto-stop, and were exercised in earlier sessions, not for this case read.
| Component | Resource | Cost |
|---|---|---|
| June chest study ingest | Server-side acquisition verb | 393 MB, 1,312 slices, 11 series; zero bytes copied to the analyst's machine |
| March to June chest comparison (rigid) | CPU, biofs-node | About 73 seconds, zero GPU |
| Native registered same-voxel A or B (right upper lobe) | CPU, rigid registration | About two to three minutes, zero GPU |
| Reconciliation reasoning workflow | LLM, five agents | 973,576 total tokens across five agents (input, output, and cached reads), about 3.4 minutes wall clock, 22 tool calls |
| Epic FHIR full import | Server-side, background | 23,949 records, about nine to twelve minutes |
| ML perception verbs (twin, findings, characterize, lesions) | A100 GPU, prior sessions | Minutes each, idle auto-stop after |
The session as a whole was an interactive collaboration spanning many turns over several days of wall-clock time, the bulk of it spent not on computation but on reasoning and on honest back-and-forth about what the numbers meant. That is the real cost profile of this kind of work, and it is also the point: the expensive, valuable resource was judgment, not floating-point operations.
5. What the orchestration took over, and what it did not
It is tempting, after a session like this, to conclude that a frontier model directly reading images will replace the specialized tools and their measurability limits. That conclusion is both wrong and, in a medical setting, unsafe. The honest decomposition is more useful.
5.1 Where the value moved
The reasoning and adjudication layer is where the value moved, and it is real. The specialized pipeline produced a confident, specific, wrong answer with a clean number attached. The orchestration caught it by pulling the organ attribution, reading the motion burden and the registration quality, and concluding that the foci were physiology, not disease. In this case, much of what the specialized pipeline contributed beyond its raw perception output was integration code and a confidence display the underlying model had not earned. That layer, the glue and the unearned confidence, is what the orchestration absorbed here. We do not claim this generalizes to all imaging tools; it is one case. The most valuable output of this case was a negative result delivered with calibrated uncertainty: the alarming number is noise, the real finding is small and indeterminate, and the next move belongs to the human and a follow-up scan.
5.2 What it does not replace
The model did not replace the perception primitives. The registration that made the same-voxel comparison possible, and the whole-body segmentation that identified the foci as motion artifact, are specialized tools the model called. When the model tried to be the perception engine, its region-focused alignment got worse, not better. And it did not replace the human. A trained radiologist saw a 3 mm shadow that the model, looking at the same native pixels, could not. The physical limit binds the model exactly as hard as it binds any convolutional detector.
The pattern. The frontier model is the attending, not the instruments and not the specialist. It orchestrates the data, calls the real perception tools as instruments, reconciles their conflicting outputs, knows the standards and the limits, and communicates with calibrated uncertainty instead of false precision. In this case it absorbed the integration code, the hand-rolled scoring heuristics, and most of all the unearned confidence. It did not absorb the segmentation and registration mathematics, and it did not absorb the expert who read a shadow it could not. The headline is not that the model replaced the tools. It is that the orchestration layer was where the conflicting signals got reconciled, the alarming number was called as noise, and the real finding was routed to the human.
6. Limitations
This is a build report and a single contested case, not a validation study. The reconciliation is a strong multi-axis argument, not a proven ground truth; the ground truth here is the human radiologist on the source images and the patient's oncology team within the trial protocol. The rigid-only registration preserves focal changes by design but cannot correct breathing, which limits the same-voxel difference at small scales. Automated detection is bounded to the detector's trained classes, so a silent automated read is not a comprehensive rule-out. The orchestration's judgment, finally, is only as good as the structured facts the tools expose to it; where a tool is wrong in a way its own outputs do not reveal, the orchestration can be led astray with it. Consent for the index case is revocable under GDPR Article 17, and the analysis is governed by the patient's Metamorphic Consent terms.
References and tools
- Eisenhauer EA, Therasse P, Bogaerts J, et al. New response evaluation criteria in solid tumours: revised RECIST guideline (version 1.1). European Journal of Cancer, 2009.
- Wasserthal J, et al. TotalSegmentator: robust segmentation of anatomical structures in CT images. Radiology: Artificial Intelligence, 2023.
- Lowekamp BC, Chen DT, Ibanez L, Blezek D. The design of SimpleITK. Frontiers in Neuroinformatics, 2013.
- NVIDIA MONAI and VISTA-3D foundation model for 3D medical image segmentation.
- MIC-DKFZ. LesionLocator: zero-shot universal tumor segmentation and tracking in 3D whole-body imaging. CVPR, 2025.
- Stanford. Merlin: a vision-language foundation model for 3D computed tomography.
- Google. MedGemma medical vision-language models.
- Anthropic Model Context Protocol specification.
- GenoBank.io. BioFS Protocol: Blockchain-Based Genomic Data Federation with DNA Fingerprints. GenoBank Whitepapers, 2025.
- GenoBank.io. GenoClaw: Patient-Owned AI Health Agents on Decentralized Genomic Infrastructure. GenoBank Whitepapers, 2026.