SpaceXAI published a post on its news site titled “Biosecurity at the frontier,” dated 1 September 2026. The post reports results from an independent LatchBio analysis of Grok 4.6.
Today, LatchBio published an independent analysis of Grok 4.6 on their biological capability and biological red-teaming benchmark suites.
What was announced
- LatchBio evaluated Grok 4.6 on BioSecBench-Refusal and BioSecBench-Surveillance.
- On BioSecBench-Refusal, Grok 4.6 was the strongest model tested at refusing disguised and hazardous tasks while still completing routine biological work. It was the only system to score above 50% on both measures.
- Reported figures include refusal of 59.2% of red-team tasks and completion of 64.8% of routine ones.
- On BioSecBench-Surveillance the model averaged a 53.5% success rate.
- SpaceXAI states it observed material improvement in Grok 4.6 biological capabilities and safeguards relative to earlier versions and outlines layered safeguards (refusal training, inference-time controls, post-deployment monitoring).
Context
The post frames the evaluation as complementary to SpaceXAI’s internal pre- and post-deployment testing. It links to LatchBio’s methods and scores at benchmarks.bio and references the Grok 4.6 model card and SpaceXAI’s Frontier Artificial Intelligence Framework.
Limits of this report
This summary covers only the statements in the official SpaceXAI news post. It does not independently verify the benchmark methodology, scores, or comparisons, nor does it assess real-world deployment performance.