The Attending Pattern: A Frontier LLM Orchestrating Specialized Tools on Patient-Owned Longitudinal CT

Daniel Uribe, founder and CEO, GenoBank.io
GenoBank Research Team
June 26, 2026

A build report. What was constructed, what it cost in time, tokens, and GPU, and where each layer reached its honest limit.

Processing identity (operator biowallet, EIP-55): 0x5f5a60Ea…d19a. The data subject in the case study is referenced pseudonymously as the index case; it is not the author.

We describe a patient-owned longitudinal CT analysis system in which a frontier large language model (Claude Opus 4.8) acts as an orchestration and adjudication layer over a set of specialized imaging tools, every one invoked as a verb of the biofs protocol against a consent-gated, on-chain-anchored imaging vault. We report the tools built (nine biofs imaging verbs and a twelve-tool imaging-radiology MCP server), the clinical-records integration that supplies the radiologist ground truth (Epic FHIR R4), and an index oncology case in which the automated change-detection pipeline produced a confident but incorrect progression call. The orchestration layer correctly identified that call as a respiratory-motion and registration artifact, across five independent lines of evidence. We are deliberate about the limit. The clinically relevant finding, a 3 mm sub-measurable nodule, sat below the resolution of every automated tool and below the model's own direct reading of the pixels. It was found by a human radiologist, and its disposition rests on short-interval follow-up imaging. The contribution is not autonomous diagnosis. It is a reproducible, provenance-anchored pipeline plus an orchestration layer whose value is calibrated judgment: knowing when an automated number is meaningless, reconciling conflicting signals, and deferring to the human expert where physics and the trained eye still win.

1. Motivation

A cancer patient on a clinical trial accumulates serial CT studies over months. The clinically important question is rarely "what is in this scan" in isolation. It is "what changed since last time, and does it matter." Answering it well requires registration across timepoints, change detection, measurement under a response standard, and, above all, judgment about which numbers are trustworthy. Automated imaging tools are improving fast at the first three. They remain weak, and frequently overconfident, at the fourth.

The motivating event for this report is concrete. An internal change-detection pipeline, run on a patient-owned vault, flagged a new lung lesion of roughly 12 mm and issued a RECIST call of progressive disease on a three-month chest interval. A reading radiologist on the same study described instead a new 3 mm ill-defined nodule, favored as a small focus of infection, with no progressively growing suspicious nodule. One of these reads, if acted on, would change a treatment plan. They cannot both be right. Resolving the discrepancy honestly, and understanding why the automated tool was confident and wrong, is the subject of this paper.

The system is built on a simple ownership premise that GenoBank.io has carried across genomics and now extends to imaging: the patient holds the asset. Each imaging study is acquired into a vault the patient controls, registered as a content-addressed biological object (a BioCID) with on-chain anchoring and revocable consent, and never copied to an analyst's laptop. The heavy bytes stay server-side; only small derived results travel.

2. Architecture

Every operation, acquisition, registration, segmentation, change detection, response measurement, and erasure, is a verb of the biofs protocol ("biological filesystem"). A verb is a thin client that submits a job manifest to the biofs-node orchestrator. The orchestrator is the only component that mounts cloud storage, schedules GPU, anchors provenance on-chain, and assembles a result manifest. This keeps the system reproducible: a result is reproducible by typing a verb, not by re-running a one-off script.

The LLM does not replace this layer. It calls it. The LLM is the orchestrator and reader of the orchestrator's outputs, the layer we describe as the attending.

2.1 The biofs imaging verbs

VerbFunctionCompute
imaging acquire (eUnity sync)Pull a DICOM study from a hospital viewer into the patient vault, server-side, deduplicated by Study Instance UID, anchored as a BioCID.Server-side I/O
imaging compareImage Time Machine. Rigid, deformation-free slice-by-slice comparison of two timepoints, with blink, crossfade, and signed-difference views and an optional z-band region crop. Rigid only, so a real focal change is never warped away.CPU
imaging twinN-timepoint registered organ meshes (a strict-isometry digital twin) for volume trends across studies.GPU
imaging findingsTumor detection and longitudinal lesion tracking with a response read, bounded to the detector's classes.GPU
imaging characterizePer-candidate feature panel (engineered radiomics and several learned backbones) plus a vision-language read and a lesion-versus-artifact head.GPU
imaging lesions (volumetric RECIST)Independent per-timepoint segmentation, cross-timepoint linking, true RECIST category, and volume doubling time, to catch iso-attenuating growth a density subtraction misses.GPU
imaging trajectoryPer-voxel persistent-versus-transient classifier across three timepoints, to separate a monotonic change from a one-interval blip.CPU
imaging attributeFrame-exact organ labeling of each detected focus by whole-body segmentation, so every change carries an organ tag and a confidence.GPU
imaging enrichFoundation-model overlays (a CT embedding model, a medical vision-language model, and a tumor segmenter) attached to a comparison.GPU

2.2 The imaging-radiology MCP server

A Model Context Protocol server exposes the orchestrator's outputs to the LLM as read tools, so the model reasons over structured facts rather than free narration. It carries the standards it applies (RECIST 1.1 measurability rules, response categories, and the ACR and RSNA structured-report sections) with citations, so any read is transparent and grounded. Tools include the comparison listing and the raw change analysis, a deterministic RECIST read and a per-lesion volumetric RECIST read, an organ-attribution read, a longitudinal pathogenic-findings read, an ACR and RSNA structured report assembler, the applied-standards reference, and four literature tools for surveying the field. Every output is framed as decision support, not a diagnosis, and confidence is graded and reported rather than withheld.

2.3 The clinical-records spine (Epic FHIR R4)

Imaging without the clinical record is half a picture. The same vault holds the patient's Epic records, pulled through a SMART-on-FHIR integration the patient authorizes. During this work we repaired four real defects in that integration and verified the repairs end to end:

The fourth defect was a dashboard button that started the authorization flow. It silently returned when no hospital was selected from a dropdown that defaulted to an empty placeholder, so a click produced nothing and reached no server endpoint. The fix auto-selects the patient's primary hospital and falls back to the first real option, so the flow is a single click.

3. Case study: a contested chest interval

We applied the pipeline to an index oncology case, a patient with pancreatic adenocarcinoma after a Whipple resection, on a clinical trial, with serial chest and abdomen CT. The data subject is consented and pseudonymous. The contest is the one described in the motivation: an automated read of new 12 mm lesion and progressive disease, against a radiologist read of a new 3 mm indeterminate nodule and no progression.

3.1 Reconciliation

An orchestrated workflow of five independent reasoning agents, each grounded in the pipeline's actual job outputs, converged on a single verdict: the automated 12 mm finding and the radiologist's 3 mm nodule are different objects, not a mis-sizing of the same object. The evidence was convergent, with several axes independently disqualifying.

AxisRadiologist 3 mm nodulePipeline flagged foci
LocationPosterior right upper lobe, high thoraxLung base near the diaphragm, opposite end of the field of view
Organ attributionLung parenchymaUnsegmented or bone, stomach, aorta, and colon at the field-of-view edge, confidence near zero, none in lung
SizeAbout 3 mm, roughly 0.014 mLTens of mL, a thousandfold mismatch
MechanismDiscrete focusLarge appeared blobs paired with co-located resolved blobs at the same slices, the signature of tissue shifting between scans
Engine self-gradeNot applicableProgression signal at low confidence, high regional motion burden near 986 mL, lowest bone-overlap quality of any chest pair

The 12 mm in the automated headline was the long axis its lesion-tracking path reported for the new lesion. The separate change-detection job surfaced the same lung-base motion as much larger appeared blobs, on the order of 30 to 50 mm long axis. Both are the one artifact seen at two processing stages. The true 3 mm nodule, meanwhile, was never surfaced by the change job at all. At roughly 0.014 mL it sits more than twentyfold below that job's 0.3 mL reporting floor, and under its downsampled detection grid. Its absence is expected at that size and is not evidence against the radiologist. The automated progression call was a registration and respiratory-motion artifact at the deformable lung base, and should not have reached a clinician as progression.

3.2 Looking at the pixels directly

We then did the thing that, on its face, sounds like it should make the specialized tools unnecessary: the model rendered the exact slices from the vault DICOM and looked at them. Series 4, the 0.6 mm lung-kernel reconstruction, at the cited image and its neighbors. Several window settings. The March study beside the June study. Finally a native-resolution registered comparison: the model rigidly aligned the full March volume into the June grid (correcting an absolute table-position offset of roughly 200 mm) and rendered a same-voxel difference at the nodule level.

The honest result. The model could not confirm the 3 mm nodule. In the registered difference the bright structures were vessel-edge residual from breathing, which rigid registration cannot undo, and which at the 3 mm scale both masks and mimics a real focus. What the model could do was exclude any gross or large new lesion, which independently corroborated that the 12 mm progression call was not a real mass, and note that a mediastinal window showed no solid soft-tissue dot, weakly consistent with the radiologist's lean toward an infectious, non-solid focus. The model was bound by the same physical limit as the tools. A 3 mm nodule is sub-measurable under the response standard, near the partial-volume floor of the scanner, and below the reliable detection threshold of every automated detector in the stack. Respecting that limit honestly is different from beating it. The radiologist found the nodule. The model did not.

The abdomen and pelvis study from the same date, read by the institution's radiologist, showed no evidence of recurrent disease after the Whipple resection, no lymphadenopathy, and no ascites. The combined restaging picture was therefore reassuring, with the one indeterminate finding a tiny nodule favored as infectious, to be settled by a short-interval follow-up CT.

4. Cost

We report cost honestly, because the most useful single fact here is an efficiency one. The entire reconciliation and direct-imaging analysis used no GPU. The comparison and the native registration are CPU jobs on commodity server hardware. The GPU-dependent perception verbs (the digital twin, detection, characterization, and volumetric RECIST) run on an A100 with an idle auto-stop, and were exercised in earlier sessions, not for this case read.

ComponentResourceCost
June chest study ingestServer-side acquisition verb393 MB, 1,312 slices, 11 series; zero bytes copied to the analyst's machine
March to June chest comparison (rigid)CPU, biofs-nodeAbout 73 seconds, zero GPU
Native registered same-voxel A or B (right upper lobe)CPU, rigid registrationAbout two to three minutes, zero GPU
Reconciliation reasoning workflowLLM, five agents973,576 total tokens across five agents (input, output, and cached reads), about 3.4 minutes wall clock, 22 tool calls
Epic FHIR full importServer-side, background23,949 records, about nine to twelve minutes
ML perception verbs (twin, findings, characterize, lesions)A100 GPU, prior sessionsMinutes each, idle auto-stop after

The session as a whole was an interactive collaboration spanning many turns over several days of wall-clock time, the bulk of it spent not on computation but on reasoning and on honest back-and-forth about what the numbers meant. That is the real cost profile of this kind of work, and it is also the point: the expensive, valuable resource was judgment, not floating-point operations.

5. What the orchestration took over, and what it did not

It is tempting, after a session like this, to conclude that a frontier model directly reading images will replace the specialized tools and their measurability limits. That conclusion is both wrong and, in a medical setting, unsafe. The honest decomposition is more useful.

5.1 Where the value moved

The reasoning and adjudication layer is where the value moved, and it is real. The specialized pipeline produced a confident, specific, wrong answer with a clean number attached. The orchestration caught it by pulling the organ attribution, reading the motion burden and the registration quality, and concluding that the foci were physiology, not disease. In this case, much of what the specialized pipeline contributed beyond its raw perception output was integration code and a confidence display the underlying model had not earned. That layer, the glue and the unearned confidence, is what the orchestration absorbed here. We do not claim this generalizes to all imaging tools; it is one case. The most valuable output of this case was a negative result delivered with calibrated uncertainty: the alarming number is noise, the real finding is small and indeterminate, and the next move belongs to the human and a follow-up scan.

5.2 What it does not replace

The model did not replace the perception primitives. The registration that made the same-voxel comparison possible, and the whole-body segmentation that identified the foci as motion artifact, are specialized tools the model called. When the model tried to be the perception engine, its region-focused alignment got worse, not better. And it did not replace the human. A trained radiologist saw a 3 mm shadow that the model, looking at the same native pixels, could not. The physical limit binds the model exactly as hard as it binds any convolutional detector.

The pattern. The frontier model is the attending, not the instruments and not the specialist. It orchestrates the data, calls the real perception tools as instruments, reconciles their conflicting outputs, knows the standards and the limits, and communicates with calibrated uncertainty instead of false precision. In this case it absorbed the integration code, the hand-rolled scoring heuristics, and most of all the unearned confidence. It did not absorb the segmentation and registration mathematics, and it did not absorb the expert who read a shadow it could not. The headline is not that the model replaced the tools. It is that the orchestration layer was where the conflicting signals got reconciled, the alarming number was called as noise, and the real finding was routed to the human.

6. Limitations

This is a build report and a single contested case, not a validation study. The reconciliation is a strong multi-axis argument, not a proven ground truth; the ground truth here is the human radiologist on the source images and the patient's oncology team within the trial protocol. The rigid-only registration preserves focal changes by design but cannot correct breathing, which limits the same-voxel difference at small scales. Automated detection is bounded to the detector's trained classes, so a silent automated read is not a comprehensive rule-out. The orchestration's judgment, finally, is only as good as the structured facts the tools expose to it; where a tool is wrong in a way its own outputs do not reveal, the orchestration can be led astray with it. Consent for the index case is revocable under GDPR Article 17, and the analysis is governed by the patient's Metamorphic Consent terms.

Decision support, not a diagnosis. Nothing in this paper is a medical diagnosis or a treatment recommendation. Every automated finding is a candidate for confirmation by a qualified radiologist on the source images. Categorical response calls are reported with graded confidence and are not a substitute for a reading radiologist or the treating oncology team. The load-bearing clinical action for a small indeterminate nodule is a short-interval, registered follow-up CT.

References and tools

  1. Eisenhauer EA, Therasse P, Bogaerts J, et al. New response evaluation criteria in solid tumours: revised RECIST guideline (version 1.1). European Journal of Cancer, 2009.
  2. Wasserthal J, et al. TotalSegmentator: robust segmentation of anatomical structures in CT images. Radiology: Artificial Intelligence, 2023.
  3. Lowekamp BC, Chen DT, Ibanez L, Blezek D. The design of SimpleITK. Frontiers in Neuroinformatics, 2013.
  4. NVIDIA MONAI and VISTA-3D foundation model for 3D medical image segmentation.
  5. MIC-DKFZ. LesionLocator: zero-shot universal tumor segmentation and tracking in 3D whole-body imaging. CVPR, 2025.
  6. Stanford. Merlin: a vision-language foundation model for 3D computed tomography.
  7. Google. MedGemma medical vision-language models.
  8. Anthropic Model Context Protocol specification.
  9. GenoBank.io. BioFS Protocol: Blockchain-Based Genomic Data Federation with DNA Fingerprints. GenoBank Whitepapers, 2025.
  10. GenoBank.io. GenoClaw: Patient-Owned AI Health Agents on Decentralized Genomic Infrastructure. GenoBank Whitepapers, 2026.
← Back to GenoBank.io Whitepapers