deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

OpenAI's Assistants API shutdown pushes agent builders to Responses and Conversations

OpenAI retired the Assistants API on August 26, 2026. Developers must now rebuild stateful agents on the stateless Responses and Conversations APIs, managing conversation state in their own code.

OpenAI's Assistants API shutdown pushes agent builders to Responses and Conversations

OpenAI's Assistants API reached end of life on August 26, 2026, and anything still running on it now has to move. According to a dev.to article by Alberto Montagnese, the shutdown forces a shift away from a managed, persistent-thread model toward a stateless approach built on two replacements: the Responses API and the Conversations API, with responsibility for state falling back on the developer.

What the Assistants API did

The Assistants API was OpenAI's earlier attempt to abstract away the difficulty of building conversational AI. As Montagnese explains, it provided a stateful environment built around persistent "threads", letting agents keep context across long interactions without developers having to manage conversation history themselves. It also bundled tools such as Code Interpreter and a built-in retrieval system.

The abstraction came with trade-offs. The API was often awkward to work with, its stateful design made costs unpredictable because the entire thread could be reprocessed on every turn, and performance suffered since developers had to poll for updates rather than stream them in real time.

The replacements: Responses and Conversations

The new model, centered on the Responses API, returns to a more primitive, stateless request-response flow. Persistent threads are replaced by Conversation objects that developers must create and manage, and keeping context between turns is now the application's job rather than OpenAI's.

Montagnese characterizes this as a significant architectural change, but one with payoffs: more control, better performance, and more predictable costs, since there is no hidden thread processing driving the bill. The Responses API also consolidates complex workflows into a single call, which simplifies the overall interaction model.

file_search replaces the old retrieval tool

For teams building retrieval-augmented generation, the most important change is the file_search tool inside the Responses API, which succeeds the old Retrieval tool as the standard way for models to access private documents.

The workflow is straightforward: create a vector store, upload files to it, and OpenAI's backend handles chunking, embedding, and indexing. That removes the need to build and maintain your own embedding and retrieval logic. When file_search is enabled on a call, the model decides from the user's query when to invoke it, runs a semantic search against the vector store, and works the relevant passages into its answer along with citations.

The dev.to article walks through a Python example: a conversation is created, a user message is added to it, and a response is then generated with the file_search tool configured against a specific vector store ID.

What developers gain and lose

The main gain is control. A stateless API gives developers direct authority over conversation history and state management, replacing the black box of the old persistent-thread system.

The main loss is convenience. The managed, effectively infinite-context thread is gone, so any application architected around OpenAI holding the state faces a non-trivial migration project. Notably, the article points out that OpenAI provides no automatic tool for converting old Threads into new Conversations.

There is also a technical trade-off in file_search. Handing the retrieval pipeline to OpenAI means giving up control over the chunking strategy. For highly structured or complex documents, the automated chunking may not be optimal, which can hurt retrieval quality, a limitation worth weighing against rolling your own retrieval stack.

The article draws on OpenAI's own Assistants migration guide and file search documentation.

Why it matters

This is a hard deadline that has already passed for anyone with production agents on the Assistants API, and there is no automated path from Threads to Conversations, so the migration is manual engineering work. It also signals the direction of OpenAI's platform: away from heavily abstracted, fully managed agents and toward lower-level primitives. Developers get more power and more predictable costs, but they now own state management, and anyone adopting the built-in file_search tool is trusting OpenAI's chunking decisions. Teams running document-heavy applications should evaluate whether the managed pipeline meets their retrieval quality bar before committing to it.

  • #openai
  • #api
  • #ai-agents
  • #rag
  • #developer-tools

Related posts