For most of the past several months, when someone has asked what has been taking up my time, the answer has been remarkably consistent: Agent Astro Version 3.
We have just released V3 to our existing clients, and later this month we will launch it publicly at RAPS Convergence in Charlotte, North Carolina. We will then take it to The MedTech Conference in Boston in October and MEDICA in Düsseldorf in November.
Version 3 is easily the most ambitious technology we have built. But somewhere during the process, I became increasingly convinced that one of the biggest challenges facing artificial intelligence is not intelligence at all.
It is fidelity.
Large language models are probabilistic by design. That is part of what makes them so powerful. They can interpret ambiguity, synthesize enormous volumes of information, reason across unfamiliar problems, and generate novel answers.
The same characteristic also creates a limitation. Ask a model the same question twice and it may take a different path, emphasize different evidence, or arrive at a different conclusion. For creative work, that variation can be useful. In consequential professional settings, it can be costly.
A database should not return a different customer record because a question was phrased differently. A financial system should not reconstruct yesterday’s transactions from statistical likelihood. And a regulatory intelligence system should not have to guess whether two medical devices are related when that relationship can be known.
The challenge is not to eliminate probability from AI. It is to determine what should be left to probabilistic reasoning and what should be anchored to a more stable representation of the world.

AI, Fidelity, Agent Astro Focus
Most general-purpose AI systems begin each question with a remarkable breadth of learned knowledge, but relatively little persistent understanding of the particular object, organization, or decision in front of them.
When asked about a medical device, a model may need to infer what the device is, which regulatory category it belongs to, which products preceded it, which manufacturers are involved, and which clearances, recalls, adverse events, or regulatory pathways matter. Even when the relevant documents are available, the model may be left to reconstruct the relationships among them.
That is not enough for the system we wanted to build.
Agent Astro is designed to make a probabilistic technology behave more deterministically where fidelity matters. Its intelligence does not reside only in the language model. It also resides in a persistent regulatory memory that we have spent years designing, mapping, and refining.
That memory identifies the objects that matter in medical-device regulation and preserves the relationships among them. Devices connect to manufacturers, product codes, predicates, descendants, submissions, recalls, adverse events, and regulatory histories. Evidence remains connected to the claims it supports. Project context can persist rather than being reconstructed in every conversation.
The model still performs work that benefits from probabilistic intelligence: interpreting questions, synthesizing evidence, comparing possibilities, and helping users reason through complex problems. But it does that work against a structured regulatory foundation. The system does not need to rediscover every known relationship each time it is asked a question.
The model provides reasoning and language. Astro provides a higher-fidelity regulatory state of the world against which that reasoning can operate.
More Than Retrieval
It is tempting to describe specialized AI as a general-purpose model connected to a collection of industry documents. That description misses the more difficult work.
Giving a model access to documents is useful, but access alone does not tell the system which documents are authoritative, how two records relate, whether a device is a predicate or a descendant, how regulatory lineage has developed, or which evidence should carry the greatest weight in a particular decision.
Those relationships are often more important than the individual documents.

This is why fidelity is a more useful objective than simple information retrieval. A high-fidelity system should preserve the identity of the objects it is reasoning about, the relationships among them, the context in which a question is being asked, and the evidence supporting the answer.
In practical terms, that means several forms of fidelity must work together:
- Entity fidelity: The system recognizes a specific device as that device, not merely as a similar collection of words.
- Relationship fidelity: Devices, manufacturers, predicates, descendants, product codes, recalls, and submissions remain connected through defined relationships.
- Evidence fidelity: Material claims can be traced to the sources supporting them.
- Context fidelity: What is known about a device or project can persist across tasks and conversations.
- Temporal fidelity: Regulatory information is understood as having a history, because what was true at one point may not remain true indefinitely.
- Output fidelity: Structured work can be constrained by known information rather than generated entirely from inference.
Together, these capabilities change the role of the model. It remains an essential reasoning engine, but it is no longer expected to serve simultaneously as the database, memory, evidence layer, relationship map, and final authority.
Fidelity Should Be Measurable
Architecture matters only if it improves performance in the real world.
Earlier this year, we benchmarked Agent Astro against leading general-purpose AI systems on medical-device regulatory research. The tests examined tasks that appear straightforward but expose important weaknesses in systems that must reconstruct regulatory facts and relationships probabilistically.
The results were significant. Astro produced no fabricated 510(k) numbers in the benchmark and identified the complete set of relevant recall records in the recall-count test. The strongest general-purpose web-enabled model identified 73 percent.
These findings do not mean that general-purpose models lack value. We use frontier models precisely because their capabilities are extraordinary. The results point to a different conclusion: model capability and system fidelity are not the same thing.
A powerful model can still produce an incomplete or unstable result when the surrounding system does not preserve the necessary entities, relationships, evidence, and context. A specialized architecture can improve performance by reducing the number of important facts the model is required to infer.
The full methodology and findings are available in our 2026 Benchmark Report.

Better AI Has Required More Human Expertise
One of the most interesting lessons from building Agent Astro is that more specialized AI has required more human expertise, not less.
Our team has grown to 12 people. It includes AI engineers, regulatory subject-matter experts, former FDA executives, professors affiliated with Johns Hopkins and the University of Pittsburgh, and people with industry experience at Boston Scientific, Baylis Medical, and other medical-device companies.
We have also expanded our sales capabilities and brought in designers to help translate a complicated technical system into a clear brand and usable experience.
Each part of that team contributes something the model cannot supply on its own.
Engineers build the architecture and connect rapidly changing AI capabilities to a stable system.
Regulatory experts identify the distinctions and relationships that influence real decisions. Former regulators bring an understanding of how evidence is evaluated.
Industry leaders know where regulatory work becomes difficult inside an operating company.
Designers make complex intelligence accessible. Customers expose the assumptions that do not survive contact with actual workflows.
This collaboration is central to the product. Domain knowledge cannot simply be uploaded as a set of documents. It must be expressed through the structure of the system: what it remembers, how it connects information, which evidence it prioritizes, what it verifies, and where it requires human judgment.
What Version 3 Represents
Version 3 is the clearest expression yet of this approach.
It brings persistent device and project context together with mapped regulatory relationships, traceable evidence, device lineage, recall and adverse-event intelligence, pathway analysis, predicate investigation, and tools that help professionals compare regulatory options. It can use different models for different tasks while preserving the knowledge, architecture, and workflows surrounding them.
The objective is not to make regulatory decisions automatic. Medical-device regulation involves evidence, interpretation, uncertainty, and professional judgment. Important decisions cannot be reduced to a rigid set of deterministic rules.
The objective is to make the foundation for those decisions more complete, consistent, and verifiable.
That distinction matters. The right response to probabilistic AI is not to pretend uncertainty can be removed from complex work. It is to stop asking the model to guess the things the system can know, preserve, or verify.
The Larger Direction of AI
The experience of building Version 3 has reinforced a broader view of where enterprise AI is heading.
Frontier models will continue to become more capable, and access to those capabilities will become increasingly widespread. As that happens, enduring value will move toward the systems built around the models: proprietary knowledge, structured memory, mapped relationships, evidence, orchestration, workflows, governance, and user experience.
Different applications will require different balances between probabilistic intelligence and deterministic structure. A creative tool should preserve room for surprise. A scientific, financial, legal, medical, or regulatory system should place much tighter boundaries around facts and relationships that can be established.
The most valuable AI systems will understand that distinction.
We are now beginning the next stage of that work. Existing Agent Astro clients are using Version 3, and we will introduce it publicly in Charlotte later this month before bringing it to Boston and Düsseldorf.
I expect the models beneath Astro to keep changing. I expect their reasoning, speed, and capabilities to improve substantially. The architecture around them must be built to benefit from that progress without surrendering the knowledge, context, and fidelity the application requires.
The future of AI will certainly include systems that know more. Increasingly, however, the systems we trust may be distinguished by something else: their ability to know what should be reasoned about, what should be remembered, what should be verified, and what should never have been left to probability in the first place.
The future may not belong to the AI systems that know the most. It will belong to the systems that know what must not be guessed.