· via TechCrunch
Baseten, Hugging Face and Goodfire team up on safety standards for open-weight models
Baseten's Base Labs research arm is partnering with Hugging Face and Goodfire AI on safety evaluation and monitoring for open-weight models, as thousands of safeguard-stripped abliterated variants circulate.

Baseten has launched a new safety initiative for open-weight AI models, teaming up with Hugging Face and Goodfire AI to build evaluation and monitoring infrastructure that the companies hope will become a shared standard. According to TechCrunch, the announcement was made on Wednesday alongside Base Labs, the research group Baseten spun up earlier this year.
How the partnership is framed
Base Labs will develop and publish methods for training and monitoring open models, with the aim of establishing a transparent standard that is woven into how models are trained and deployed rather than patched on after release. The companies have not yet explained how the collaboration will work in technical terms, TechCrunch reports.
Baseten laid out its reasoning in a post on X: "We believe openness to be an advantage for AI safety." The company argued that openness provides more visibility into model behavior and better ways to turn safety research into transparent, actionable controls than closed-source alternatives.
Goodfire framed the objective in a reply to Baseten's post: "Safety must be built into open models and provided by those who serve them." The startup specializes in interpretability — techniques for explaining how models reach their decisions — which TechCrunch identifies as the most likely contributor to the "built in" side of the effort.
The trio covers the open-weight stack between them: Hugging Face hosts the ecosystem of downloadable models, Baseten provides the inference infrastructure those models run on, and Goodfire supplies interpretability tooling. Baseten has also issued an open call for the wider developer community to contribute, saying the goal is "an ecosystem of open models that are safe and accessible to all."
Why abliteration is driving this
The initiative arrives amid a running debate over whether open-weight models are safe to distribute at all. The sharpest evidence is a technique called abliteration, which removes a model's built-in safeguards after the fact. Hugging Face currently lists more than 6,000 abliterated models on its platform, according to TechCrunch — a figure that illustrates how easily post-training guardrails can be undone once weights are public.
That is the problem the partnership is positioning itself against. Safeguards embedded in the model itself, and monitored during deployment by the infrastructure serving it, are harder to strip than a refusal behavior added on after training.
Deep-pocketed participants
Both lead companies are well capitalized. Baseten, an AI inference provider, raised a $1.5 billion Series F in June, lifting its valuation to $13 billion. Goodfire raised a $150 million Series B led by B Capital earlier this year to advance its model interpretability platform.
Why it matters
The partnership stands out as a coordination effort. Rather than each company shipping its own safeguards in isolation, a model host, an inference provider and an interpretability specialist are trying to define a common standard for how open models are trained and monitored. If it is adopted, safety checks could become part of the default deployment path for open models instead of an optional extra.
The stakes are visible in the 6,000-plus abliterated models on Hugging Face: guardrails applied after training are proving easy to remove. Building safety in, and watching for it at serve time, is the alternative this group is betting on. The open questions are just as clear — no technical details or published methodology exist yet, and the standard's success will hinge on whether the broader developer ecosystem Baseten is courting actually adopts it.
- #open-weight-models
- #ai-safety
- #baseten
- #hugging-face
- #interpretability