Comparison
I Tried Eight AI Biology Agents, So You Don’t Have To
Claude Science, Rosalind Workbench, LatchBio, Benchling, DNAForge, Edison, Phylo, and Beakr gave similar answers but in very different shapes.

Given all the updates in agents for biology, I’ve decided to quickly write up my experience using their products. I am reviewing Claude Science, OpenAI Rosalind Workbench, Benchling, LatchBio, DNAForge, Edison Scientific, Phylo, and Beakr, and will be placing them on a spectrum between chatbots with tools on one end and lab notebooks with AI on the other.
I found they all produce similarly good answers to a given prompt, so I’ll be focusing on the real differentiator, which is the UX and overall integration with a biologist’s workflow. Claude Science and OpenAI Rosalind Workbench felt like chatbots with scientific tools. Benchling, LatchBio, and DNAForge felt like biology workspaces with AI. Phylo, Beakr, and Edison Scientific sat somewhere between those two models.
I found Claude Science closest to a classic AI assistant but with scientific tools, while Rosalind Workbench felt like OpenAI’s response. They both embrace their chat-thread UX, with an OpenAI employee literally calling their product a “super plugin”, making it familiar for any biologist who is also a tokenmaxxers.
Beyond the chat surface, both products offer helpful visualizations and artifacts, as well as built-in UI for things like NGS. They also offer the ability to do certain work remotely, which is helpful for computationally intensive bioinformatics procedures.
At the other extreme are products that provide an entire molecular biology workspace with integrated AI, with Benchling, LatchBio, and DNAForge falling into this category.
Starting with Benchling’s AI, the underlying intelligence they provide is fairly strong and similar to the other products. However, Benchling definitely beats Claude or Codex when the AI needs to work inside an existing electronic lab notebook and use related tools. They also use their agent to import data into the notebook and connect to other systems in your lab, integrating their AI with your system of record.
LatchBio overall focuses on bioinformatics workflows, with separate tabs for Workflows (pipelines), Pods (compute), and Data (data), but its Plots tab hosts their AI chat. It can perform bioinformatics analysis, set up and run pipelines, and even has an IGV plugin. LatchBio is ideal for jobs centering on hosted bioinformatics workflows and compute.
As for DNAForge, we also have an electronic lab notebook and an IGV-style NGS viewer. But our focus is ensuring that humans and agents work with the exact same tools in the exact same workspace, allowing for real-time observability and deterministic guardrails.
DNAForge is less like asking a chatbot for an answer and more like working with an agent inside editable biological files.
Aside from the Hosted Agents we provide, you can also use your own Claude or Codex subscriptions within our harness.
In the middle of the spectrum are Phylo, Beakr, and Edison Scientific, with a classic chat thread in the middle, file-picker on the left, and tooling drawer on the right. They feel like dedicated AI research environments with more structure than a chatbot, but they do not include full electronic lab notebooks.
Despite that, they still offer hosted files/compute and features such as a Python notebook or a knowledge graph. They also include more specialized biology UI, like a sequence viewer, directly into the product.
I view the spectrum as a bet on how central human biologists will be in the next five years. If we foresee them lightly orchestrating autonomous swarms of biology agents, then minimal products Claude Science and Rosalind Workbench, or even more structured options like Phylo, Edison and Beakr, make sense. But if we see humans continuing to drive wet-lab research with agents acting as junior researchers, then workspace products like Benchling, Latch, and DNAForge are ideal.
The AI x bio space today is as crowded as coding was in early 2025, but I hope I’ve provided some legibility. I would not choose among these products based on which one gave the best answer to a single prompt, but on what I wanted to do with the answer afterwards.