Sovereign access—Join the waitlist →
AI & Research

Agents Can Notice an Anomaly. The Constraint Is the Record.

Anthropic’s Claude agents found a previously uncharacterized reverse-transcriptase system in phage DNA. Function is still unknown. The harder problem for R&D is what happens to the hypotheses a campaign like this produces — and then discards.

Roberto Honegger
Roberto Honegger
Founder, OVAITY9 min read
Floating scientific panels connecting a DNA repeat array, a bacteriophage, protein structure, agent reports, and a laboratory assay over a faint network

The topic at a glance

  • On 23 September, Anthropic reported that Claude agents surveying reverse-transcriptase loci noticed a tandem DNA repeat array that defines a jumbo-phage system they call ART.
  • The useful result is not a CRISPR comparison. Function remains unknown, most candidate reports were discarded, and the same campaign did not recover the array when rerun.
  • OVAITY is building a scientific memory layer so computational campaigns, rejected hypotheses, and lab follow-up can stay connected as one piece of work.

Last week Anthropic reported that Claude had found a previously uncharacterized enzyme system in bacteriophage DNA. The company compared the architecture to CRISPR. That comparison is doing too much work. [1]

The finding is real, and narrower than the coverage. Agents running in an autonomous harness surveyed reverse-transcriptase loci across roughly 1.9 billion protein clusters. One of them, reading raw DNA next to a jumbo-phage reverse transcriptase, noticed a tandem repeat array that earlier genome reports had not treated as part of a system. Anthropic named the family array-associated reverse transcriptases, or ART. [1,2]

The reverse transcriptase itself had already been seen. What the agents appear to have noticed first was the arrangement around it: a bank of non-coding repeats, a dedicated partner gene, and an RT with an unusually long N-terminus. Function is still unknown. The authors say they have not shown that the enzyme is active, that the array RNAs are its substrates, or what the system does for the phage. [1,2]

The operational question is not whether an agent can spot a repeat. It is what a research organization does with a campaign that generates more candidates than a bench can test, discards most of them, and still has to remember why.

What the agents actually did

Anthropic formed a life-sciences group and a Bay Area molecular biology lab in the spring of 2026. The computational work ran in Claude Science and Claude Code, with a harness coordinating many sessions in parallel. All laboratory experiments were done by human scientists, at BSL-1 and BSL-2. [1,3]

The campaign started from a research brief: look for novel reverse-transcriptase systems based on new partner-gene associations. Worker agents planned and ran tasks. Supervisors reviewed them and opened follow-ups. Plans, results and reviews went into a shared record. The run comprised 119 tasks and 949 agent sessions over 21.5 hours of wall-clock time, using about 216 million tokens, without human intervention during the search. [2]

The numbers that followed are more informative than the runtime. The agents recovered about 198,000 RT clusters, classified them into nine classes, sampled roughly 11,000 loci, and scored 3,564 recurring neighborhood protein families as candidate partners. Seventeen families were promoted for deep dives. Fourteen were later set aside as annotation artifacts, parts of known systems, or ordinary neighborhood residents. Three were retained as previously unreported associations. Workers also flagged unusual features of individual RTs, which produced three new lineages. ART was one of them. [2]

ART did not come from the partner-gene census as originally scoped. A worker first noticed the RT because it sat near a phage RNA-polymerase subunit, then rejected that gene as a dedicated partner, then queued a follow-up. A later worker compared upstream DNA of related RTs, loaded a metagenomic flank into context, and recognized a tandem repeat. It counted copies, measured spacers, compared the layout with retrons, CRISPR arrays and diversity-generating retroelements, and searched the literature before filing a report. [2]

That sequence is closer to expert curation than to a predefined pipeline. Conventional genome-mining methods retrieve homologs and sort them by features specified in advance. Anomalies outside those features go unseen. The authors’ point is that an agent reading primary sequence can supply some of the judgment that used to sit only with the expert at the end of the pipeline. [2]

What reached the bench — and what did not

Follow-up searches found 95 distinct ART RT clusters in cultured jumbo phages and predicted viral contigs. Twenty-eight carried a detectable upstream array. Repeats are 15 to 49 nucleotides, with palindromic cores, separated by unique spacers of 120 to 220 nucleotides — much longer than typical CRISPR spacers. No cas genes sit near the loci. [2]

The biological evidence so far is transcriptional, not functional. Reanalysis of a published Staphylococcus phage SA1 infection time course showed array-derived RNAs accounting for up to 8% of phage RNA at 15 minutes post infection, among the most abundant phage transcripts. Expressing the SA1 ART locus on plasmids in Escherichia coli produced discrete short RNAs with similar boundaries. That is consistent with a retron-like architecture that carries a bank of RNAs where a retron carries one. It is not evidence of gene editing, programmability, or phage defense. [2]

Anthropic’s own language is more careful than some of the headlines. Feng Zhang, reviewing the preprint, called the identification of RNA-repeat arrays associated with reverse transcriptases “genuinely intriguing and merits further investigation.” That is the right register. [1]

Eric Kauderer-Abrams, Anthropic’s head of life sciences, told Reuters a few days before the paper that “to do biology, the final test is still and will be for a while in real lab work.” The ART announcement is consistent with that. Agents proposed. Humans ran the experiments that currently exist. The experiments that would establish function are still underway. [1,3]

A campaign that will not replay still needs a record

The authors reran the same campaign ten more times. Nearly every completed census sampled ART loci. In two runs, workers investigated the lineage. None of the reruns read the upstream DNA in a way that recovered the array. The authors attribute this to the size of the RT search space and the non-deterministic behavior of the harness. [2]

That result is easy to skip past. It is the part that matters for anyone trying to put agentic genome mining into an R&D organization.

If the discovery depends on a particular agent loading a particular flank into context, then the session transcript is part of the scientific record. Lose it, and you lose the observation. The original MarsHill genome report had already identified the RT and proposed a 5′ non-coding RNA, without describing the repeats or the partner gene. The array was sitting in public data. Noticing it once does not mean the organization can notice it again, or explain to a later scientist why this locus was promoted and a neighboring one was not. [2]

Benchmarks in the preprint sharpen the same point. Direct observation of DNA in the model’s context was critical for recognizing the array. When the sequence sat in files and had to be opened with tools, many attempts never read a contiguous stretch long enough to see more than one repeat unit. Recognition rose when enough sequence was actually read. An agent that has access to a dataset is not the same as an agent that has seen the relevant stretch. [2]

This is a more specific version of a problem we have already described for bioinformatics analyses and for multi-agent systems that reason over public biomedical databases. Results survive. The path that produced them often does not.

The surplus of hypotheses is now the operational problem

Anthropic is explicit about a second consequence. Because Claude produces hypotheses so prolifically, “the hypotheses themselves have become an object of study.” A single campaign can yield hundreds to thousands of candidate reports. The lab then has to decide which proposals are worth testing, and feed that judgment back into the instructions given to the agents. [1]

That is a knowledge-management problem wearing a discovery story.

Most of the 17 promoted partner families were rejected. Those rejections are scientifically useful: they record what looked like a novel association and was not. In a conventional lab, that information lives in someone’s head, a Slack thread, or nowhere. When the next campaign runs, an agent without that memory can spend tokens rediscovering the same artifacts.

The same is true on the experimental side. Array expression in E. coli is a result. The assays that have not yet been done — RT activity, RNA substrates, partner interaction — are also part of the project state. An agent helping to interpret the next gel needs to know which hypotheses were already retired, which constructs were already expressed, and which claims the preprint declined to make.

Earlier this month, Stanford’s Virtual Biotech showed that specialist agents can assemble inspectable cases over public resources. Anthropic’s ART campaign shows they can also notice anomalies in primary sequence. Neither result substitutes for a durable record of the organization’s own computational campaigns and bench follow-up. Claude Science can sit in the workflow. It does not, by itself, remember why a candidate was set aside six months ago.

What has to exist between the report and the next experiment

If agentic genome mining becomes a regular part of discovery groups, a few ordinary capabilities have to sit underneath it.

  • A queryable archive of campaign artifacts: research briefs, task trees, session transcripts, rejected candidates, and the reports that reached a human.
  • Provenance that ties a wet-lab result back to the computational claim it was meant to test, including the sequence version, construct, and the reason that particular locus was chosen.
  • A way for a scientist to inspect why an agent loaded one flank and not another, especially when the same campaign does not replay.
  • Scoped permissions. Public metagenomic clusters are a different governance problem from an organization’s unpublished constructs, infection data, and partner-gene hypotheses.

The FAIR principles were written for reuse by humans and machines. Agent campaigns make machine-actionability a weekly requirement rather than a repository afterthought. [4]

This is also one of the challenges we are exploring at OVAITY. We are building a Scientific Intelligence Platform that connects projects, experiments, datasets, analyses and decisions in one governed environment, so scientific context remains usable over time. Context-aware AI over that project record is being validated with researchers. Broader research agents and autonomous agent workflows remain prototypes. We are not building a 950-agent genome-mining harness. We are trying to make the underlying scientific record coherent enough that a computational campaign, the hypotheses it discarded, and the experiments that followed can be inspected as one piece of work.

Treat the campaign as a scientific object

The ART family may yet turn out to be an interesting phage system. It may not. That uncertainty is normal, and the authors have been clear about it.

The workflow, though, is already here. Agents can search sequence space faster than expert curation, notice features a predefined pipeline would miss, and write reports faster than a bench can test them. Keep the transcripts, the discarded candidates, and the bench outcomes in one place, and the next campaign starts from a scientific position rather than from a blank prompt.

Until that record exists, adding more agents mainly produces more reports about public DNA.

If your team is exploring AI-native research workflows but struggling with context, traceability, or project memory, we’d love to learn from you.

References

Sources

  1. Yoon, P. H., Athukoralage, J. S., Ameisen, E., Kauderer-Abrams, E., Perry, N. T. and Durrant, M. G.

    Autonomous AI agents discover reverse transcriptases with tandem repeat arrays

    Anthropic technical report · Sep 23, 2026

  2. Dastin, J. and Erman, M. / Reuters

    EXCLUSIVE: Anthropic quietly sets up biology lab as it ramps AI drug program

    Reuters · Sep 18, 2026

  3. Wilkinson, M. D., Dumontier, M., Aalbersberg, I. J. et al.

    The FAIR Guiding Principles for scientific data management and stewardship

    Scientific Data · 2016