Sign in

SecureBio, Inc.

@securebio.org
94 followers 1 following 145 posts
PostsRepliesMedia
SecureBio, Inc. @securebio.org · 22/09/2026
Frontier AI now beats the best human virologists on biosecurity-relevant evals. Once an AI model bests every human expert we can test it against, a higher score stops telling us anything about real-world risk. Read our new blog post on how bio evals have to adapt: securebio.org/blog/evaluat...
Diagram showing that as AI model capabilities move from beginner to expert to superhuman, evaluations for chess and cybersecurity can shift from human-built proxies (puzzles, coding benchmarks) to real endpoints (real games, real exploits), while biosecurity cannot safely make that shift and must instead rely on indirect signals: experimental prediction, design tasks, defensive acceleration, and adjacent domains.
021
SecureBio, Inc. @securebio.org · 18/09/2026
Untargeted wastewater metagenomic sequencing screens hundreds of pathogens at once, but that breadth costs coverage, making variant tracking more difficult. However, by pooling CASPER reads nationally, we can recover real lineage and genotype trends. Read the latest from SecureBio Detection here ⬇️
securebio.org
Tracking pathogen variants with untargeted wastewater metagenomics – SecureBio
000
SecureBio, Inc. @securebio.org · 17/09/2026
Serious about AI safety? SecureBio AI conducts pre-release assessments and independent risk evaluations for multiple frontier models. We are hiring for engineers, researchers, and many other roles to expand this work. Reach out if you're interested in working on bio-focused third-party assessment!
securebio.org
Work with us – SecureBio
000
SecureBio, Inc. @securebio.org · 14/09/2026
Ten hours of manual testing yielded similar findings, showing generally robust refusal behavior but some gaps, and a willingness to help with biological tool use and engage in high-level discussions. Our full report is here: securebio.org/resources/gp...
securebio.org
000
SecureBio, Inc. @securebio.org · 14/09/2026
With production safeguards on, GPT-6 Astra refused or blocked 82.9% of hazardous BioTIER queries (vs. 66% for GPT-5.6 Sol) while still complying with 95.3% of benign ones, and refused the screening-evasion task 100% of the time.
100
SecureBio, Inc. @securebio.org · 14/09/2026
On our agentic benchmarks, GPT-6 Astra modestly outperformed GPT-5.6 Sol on ReproBAIT and was the first model we've tested to successfully outline a complex, known DNA screening evasion strategy for ASE.
100
SecureBio, Inc. @securebio.org · 14/09/2026
On our knowledge benchmarks, Astra exceeded every previously tested model on VCT and VCT-v2 and performed within range of GPT-5.6 Sol on HPCT. On WCB, it scored slightly lower than earlier models, driven largely by increased hedging.
100
SecureBio, Inc. @securebio.org · 14/09/2026
We ran a brief pre-release assessment of OpenAI's GPT-6 Astra that included manual testing and key benchmarks evaluating misuse-relevant knowledge, agentic capabilities, and refusal behavior for hazardous biology queries. Read our summary here: securebio.substack.com/p/securebio-...
securebio.substack.com
SecureBio’s GPT-6 Astra Pre-Release Testing Report
We report findings from our pre-deployment assessment of the biological and biosecurity-relevant capabilities of GPT-6 Astra.
130
SecureBio, Inc. @securebio.org · 11/09/2026
Model accuracy dropped only 1.0–7.1 points on VCT-v2, and the gap between strongest and weakest models held steady, so past VCT results remain largely valid. VCT-v2 also has more headroom before saturation: at least ~30 points vs. ~23 on the original VCT.
Heatmap showing per-question performances of models and human experts on VCT-v2. Each cell represents the mean question × model accuracy, with darker colors indicating higher accuracy. Mean model accuracy on all questions is displayed to the right of the heatmap. Yellow cells represent data missing due to model refusals or missing expert baseline entries. The theoretical best model is taken to be the top-performing model on a per-question basis.
000
SecureBio, Inc. @securebio.org · 11/09/2026
We found that 27% of the original VCT questions were shortcut-exploitable (answerable without the image or question text). In all, we edited 163 questions (50.6%) and removed 43 questions (including 25 non-discriminating ones) to produce the 279-question VCT-v2 benchmark.
Flowchart showing the generation of the updated 279-question VCT-v2. Of the 322 VCT questions, 43 were removed, including 25 easy, non-discriminating questions; 163 were edited to improve their scientific accuracy, clarity, and/or resistance to heuristic shortcuts; and 116 remained unchanged.
100
SecureBio, Inc. @securebio.org · 11/09/2026
Top model accuracy on the original VCT has risen ~12 points since April 2025 (o3: 42.0%; GPT-5.6 Sol: 53.9%). To ensure this benchmark continues to capture the full dynamic range of model capabilities, we audited all 322 questions for validity and shortcut-resistance.
Benchmarks are restricted by the reliability limit of the answer key. This sets the “effective ceiling” of benchmarks and causes benchmark accuracy to stop scaling linearly with the scientific capability of the model beyond a certain threshold, resulting in lost dynamic range. Note: This schematic is an illustrative representation only.
100
SecureBio, Inc. @securebio.org · 11/09/2026
We updated our flagship virology troubleshooting benchmark, the Virology Capabilities Test (VCT), to be more shortcut-resistant and to have a larger headroom. We confirmed that the majority of VCT questions are scientifically valid and that neither VCT nor VCT-v2 is saturated.
securebio.substack.com
Introducing VCT-v2 — the updated Virology Capabilities Test
Updating Virology Capabilities Test (VCT) to be more shortcut-resistant and to stay ahead of saturation.
111
SecureBio, Inc. @securebio.org · 20/08/2026
We've shown that pathogen-agnostic detection works; now's the time to show it can be done fast enough to provide meaningful protection.
010
SecureBio, Inc. @securebio.org · 20/08/2026
With support from Coefficient Giving and others, we've collaborated with our CASPER partners to massively scale up wastewater sequencing, expanded to nasal swab sequencing, and built pipelines for detecting engineered sequences.
100
SecureBio, Inc. @securebio.org · 20/08/2026
We view pathogen-agnostic metagenomic sequencing as the best opportunity for identifying novel pandemic threats, going beyond lists of known infectious agents to capture the wide range of nucleic acids in a sample.
100
SecureBio, Inc. @securebio.org · 20/08/2026
The OpenAI Foundation has granted SecureBio Detection $17.2M to reduce our end-to-end time (sample collection to results) from 14 to 3 days, expand our collection footprint, and further validate our detection system: securebio.org/blog/three-d...
securebio.org
Building a three-day early-warning system for novel pathogens – SecureBio
111
SecureBio, Inc. @securebio.org · 12/08/2026
Finally: We're hiring! Open roles in lab science, software engineering, project management, logistics, IT, and public health response, with referral bonuses of $2k-$6k. securebio.org/careers/
securebio.org
Work with us – SecureBio
000
SecureBio, Inc. @securebio.org · 12/08/2026
We are continuing to refine our processes and methods for escalating hits on our biosurveillance and notification system, SecureBio Alerts. We cover several of these (and some other particularly interesting CASPER detections) in a recent blog post. securebio.org/blog/rare-vi...
securebio.org
Rare viruses, vaccines, and more in over 1 trillion metagenomic reads – SecureBio
100
SecureBio, Inc. @securebio.org · 12/08/2026
We’ve also been busy with new collaborations: expanding Zephyr to Miami with Helena Solo-Gabriele, a 10-week pilot on transmission dynamics in mass gatherings with Hannah Healy and the BPHC, and testing out GPT-Rosalind.
100
SecureBio, Inc. @securebio.org · 12/08/2026
On the computation side, we’ve completed initiatives to support day-to-day lab work, made substantial progress in data pipeline automation, and worked to incorporate frontier AI, developing standardized sandbox instances to safely run AI agents with minimal restrictions.
110
SecureBio, Inc. @securebio.org · 12/08/2026
We completed our move to our own lab space in Kendall Square in May. The larger footprint has let us increase our sample throughput and add new equipment (like our new MiSeq i100 Plus), and the proximity to the office makes it easier for lab and non-lab staff to coordinate.
100
SecureBio, Inc. @securebio.org · 12/08/2026
To make the Zephyr data as useful as possible, we’ve shared all 179 million human-scrubbed ONT longreads generated through June 24, 2026 (from 25,790 unique nasal swabs) on SRA. We encourage researchers to use this dataset and contact us with any questions. www.ncbi.nlm.nih.gov/bioproject/P...
ncbi.nlm.nih.gov
Pooled respiratory samples (ID 1379685) - BioProject - NCBI
A BioProject is a collection of biological data related to a single initiative, originating from a single organization or from a consortium. A BioProject record provides users a single place to find l...
100
SecureBio, Inc. @securebio.org · 12/08/2026
We made several changes to Zephyr sample processing that significantly improved viral yields. The resulting improved ability to assemble complete viral genomes will be important for downstream work (e.g, test design, vaccine/antiviral development) for novel potential pathogens.
100
SecureBio, Inc. @securebio.org · 12/08/2026
Meanwhile, our nasal swab sampling program, Zephyr, achieved a new single-day collection record (401 swabs) on May 9. Real-time data from this work is available on our Zephyr Dashboard. data.securebio.org/zephyr/
Screenshot of the Zephyr dashboard showing a bar graph with the number of samples collected per day, highlighting May 9th (401 samples).
100
SecureBio, Inc. @securebio.org · 12/08/2026
Lenni Justen (Sabeti lab) also published a preprint on normalization for wastewater metagenomic sequencing for quantitative pathogen tracking, identifying ways to make these data even more effective for monitoring trends in viral abundance and spread. www.medrxiv.org/content/10.6...
From Justen et al. (2026): Wastewater metagenomic sequencing (WW-MGS) normalization marker abundance and dynamics. (b) A dot plot showing per-site median marker fractions; each point represents one site (n = 25), the horizontal black bar marks the across-site median, and the colors represent different geographic sites. (c) Pairwise within-site Pearson R on log10-transformed marker fractions. Each site's mean log10(marker) is subtracted before pooling, so correlations reflect sample-to-sample co-variation within sites. 

Establishing wastewater metagenomics as a quantitative
pathogen monitoring tool with normalization
Lennart Justen, Alessandro Zulli, Rose S. Kantor, Rebecca Y. Linfield, Leon S. Moskatel, Daniel Cunningham-Bryant, Jeff Kaufman, Marc C. Johnson, Michael R. McLaren, Pardis C. Sabeti
doi: https://doi.org/10.64898/2026.07.14.26356442
https://www.medrxiv.org/content/10.64898/2026.07.14.26356442v1
100
SecureBio, Inc. @securebio.org · 12/08/2026
Data from CASPER wastewater surveillance continue to yield key insights, like this wastewater RNA virome analysis from Rose Kantor (LLNL), Migun Shakya (LANL) et al. that may increase the efficiency and accuracy of future human pathogen screening efforts. www.medrxiv.org/content/10.6...
From Kantor et al. (2026): A bar graph showing the relative abundance of vOTUs from different viral orders in the wastewater virus genome database (WVDB). Bars are colored by the predicted host for each vOTU, and “other” includes vOTUs classified to families for which the ICTV virus properties table lists multiple hosts. Hosts are shown as “unknown” where no order- or family-level taxonomic call was made and the vOTU had no genus-level BLASTN hit against NCBI core-nt. 

A genome-resolved view of the wastewater RNA virome
 View ORCID ProfileRose S. Kantor, Migun Shakya, Nelson Ruth, Jason A. Rothman, Clayton Rushford, Devon A. Gregory, Aidan Epstein, Jeff T. Kaufman, Jonathan E. Allen, Patrick S. G. Chain, David H. O’Connor, Marc C. Johnson
doi: https://doi.org/10.64898/2026.05.19.26353600 
https://www.medrxiv.org/content/10.64898/2026.05.19.26353600v2
100
SecureBio, Inc. @securebio.org · 12/08/2026
SecureBio Detection has been busy! We’ve expanded our biosurveillance network, made our nasal swab program more sensitive and built new partnerships, bringing us closer to a system that can rapidly detect and report on unusual biological events. Get a full rundown here: securebio.org/blog/updates...
securebio.org
SecureBio Detection Updates, August 2026 – SecureBio
110
SecureBio, Inc. @securebio.org · 28/07/2026
We conservatively apply these conclusions to the later Opus 4 models but not Fable 5, Mythos 5 or Opus 5. Read the full review here: Summary: securebio.substack.com/p/review-of-... Full review: securebio.org/resources/an...
securebio.substack.com
Review of Anthropic’s Unredacted Chemical and Biological Risk Report: Claude Opus 4.6
We share our external review of the chemical and biological section of Anthropic's February 2026 Risk Report after months of followup and hundreds of pages reviewed.
000
SecureBio, Inc. @securebio.org · 28/07/2026
Fact 3: We emphasize that risks remain non-negligible. Actors with extremely high jailbreaking expertise can still quickly find or make substantial progress on jailbreaks around the refusal classifiers (e.g. UK AISI’s BPJ) and remediation times can be long.
100
SecureBio, Inc. @securebio.org · 28/07/2026
Fact 2: We reviewed Anthropic’s process for (a) vetting users for exemptions from refusal classifiers and (b) monitoring such users. We find it unlikely but possible that a threat actor would acquire and use such an exemption to substantially uplift CB weapons production.
100
SecureBio, Inc. @securebio.org · 28/07/2026
However, we advise that the constitutions should be continually updated as advances in biotechnology and elsewhere change the landscape of topics that pose CB risk.
101
SecureBio, Inc. @securebio.org · 28/07/2026
Fact 1: We examined the training constitutions and robustness of Anthropic’s refusal classifiers, their main CB safeguard. The classifiers generally refuse on topics we would have refused on, including 94.2% of hazardous prompts in BioTIER-refuse.
100
SecureBio, Inc. @securebio.org · 28/07/2026
Overall, we agree with Anthropic that the current risk of catastrophic outcomes substantially enabled by Claude Opus 4.6 is: - “Very low but not negligible” for non-novel CB weapons production, and - “Low risk, but with substantial uncertainty” for novel CB weapons production.
100
SecureBio, Inc. @securebio.org · 28/07/2026
We also independently assessed the coverage of the Claude 4.6 and 4.7 classifier guards with BioTIER, SecureBio’s set of biohazardous prompts.
100
SecureBio, Inc. @securebio.org · 28/07/2026
Over two months, Anthropic shared 110 pages of additional materials and responded to over 70 follow-up questions regarding model capabilities and safeguards. We probed their refusal classifier training constitutions, prompts used to assess their classifiers, and uplift study methodologies.
100
SecureBio, Inc. @securebio.org · 28/07/2026
Summary: securebio.substack.com/p/review-of-... Full review: securebio.org/resources/an...
securebio.substack.com
Review of Anthropic’s Unredacted Chemical and Biological Risk Report: Claude Opus 4.6
We share our external review of the chemical and biological section of Anthropic's February 2026 Risk Report after months of followup and hundreds of pages reviewed.
100
SecureBio, Inc. @securebio.org · 28/07/2026
Anthropic collaborated with us on a review of the unredacted chemical and biological (CB) risk sections of their February 2026 Risk Report. After months of followup and hundreds of pages reviewed, we share our external review below.
securebio.substack.com
Review of Anthropic’s Unredacted Chemical and Biological Risk Report: Claude Opus 4.6
We share our external review of the chemical and biological section of Anthropic's February 2026 Risk Report after months of followup and hundreds of pages reviewed.
110
SecureBio, Inc. @securebio.org · 24/07/2026
We thank the OpenAI team for engaging with us on pre-release testing. You can read more at the link above.
000
SecureBio, Inc. @securebio.org · 24/07/2026
On our latest refusal benchmark BioTIER, GPT-5.6 Sol refused 66% of high-risk bio prompts and correctly answered 99% of safe ones. Top models refuse over 90% of high-risk prompts.
100
SecureBio, Inc. @securebio.org · 24/07/2026
On a DNA-synthesis-screening-evasion task, the safeguarded launch model refused outright, as did other leading guarded models. Only the unrestricted "railfree" variant completed it, using a known but impractical method.
100
SecureBio, Inc. @securebio.org · 24/07/2026
GPT-5.6 Sol beat every prior model on three of our four bio knowledge benchmarks and scored above our PhD-level experts on all four. On World-Class Bio it hit 68%, roughly 9 percentage points above OpenAI's previous flagship.
100
SecureBio, Inc. @securebio.org · 24/07/2026
We at SecureBio tested GPT-5.6’s biorisk-related capabilities: virology and pathogen knowledge, niche scientific knowledge, agentic bio capabilities, and bio AI tool usage. securebio.substack.com/p/gpt-56-sol...
securebio.substack.com
GPT-5.6 Sol Pre-Release Testing Report
We report findings from our comprehensive pre-deployment assessment of biological and biosecurity-relevant capabilities of GPT-5.6 Sol.
100
SecureBio, Inc. @securebio.org · 10/07/2026
(i) Continue giving access to advanced models to 3rd party orgs for testing. (ii) Keep building and standardizing the evaluations ecosystem. (iii) Prevent advanced bio capabilities from reaching unverified users by default. Link to article below: securebio.substack.com/p/preparing-...
securebio.substack.com
Preparing for the “Bio Mythos” Moment
Mythos showed that governments are underprepared for emerging dangerous AI capabilities. To prepare for advanced bio capabilities, we should build biosecurity assurance infrastructure proactively.
000
SecureBio, Inc. @securebio.org · 10/07/2026
What can we do to prepare for a "Bio Mythos" moment? SecureBio's Coleman Breen and Hodan Omaar expand on the relationship between cyber and biological capabilities of advanced AI models and provide three recommendations for policymakers.
100
SecureBio, Inc. @securebio.org · 09/07/2026
Part of the power of our sequencing approach is being able to distinguish between benign and concerning sequences. We can learn a lot from just a few nucleic acids in the sewer system and we use these signals to launch informed responses. Read the full post here: securebio.org/blog/rare-vi...
000
SecureBio, Inc. @securebio.org · 09/07/2026
Sometimes our findings require action, but we see a lot of benign things too. Our systems that identify engineered pathogens can also find vaccine constructs or lab materials that pose no risk to the general population.
110
SecureBio, Inc. @securebio.org · 09/07/2026
We do a lot of wastewater sequencing, 80 billion reads a week to be exact. In sewersheds across the country we monitor for pathogens that can cause pandemics using ultradeep, metagenomic sequencing.
100
SecureBio, Inc. @securebio.org · 18/06/2026
SecureBio is excited to contribute to the World Cup wastewater surveillance efforts led by the Health Security Operations Center.
010
SecureBio, Inc. @securebio.org · 17/06/2026
The dashboard is live and updated as new models, benchmarks, and analyses are released. Check it out here: securebio.org/benchmarks/
000
SecureBio, Inc. @securebio.org · 17/06/2026
In addition to capabilities, we also measure and report the effectiveness of biosecurity-related safeguards, such as BioTIER. All scores are computed using a standardized pipeline to ensure fair and valid comparisons. For example, we make sure evaluations are run using identical configurations.
100