Instrument theme
Technology · September 2026

Agents that carry
the work forward.

Managed agent runtimes take responsibility for the loop around a model: continuing a task, maintaining context and coordinating tools. OpenAI’s new Agents API makes this an option for teams building enterprise workflows.

The Little Builder coordinating connected document, tool and completed-work stations
Scope an agent pilot

OpenAI Agents API · Public beta · Reviewed 19 September 2026

WHAT CHANGED

The execution loop becomes a service.

OpenAI released the Agents API in public beta on 10 September 2026. It manages durable sessions, orchestration, context compaction and recovery, with hosted or connected sandbox environments. Official release changelog.

The runtime can execute code, work with files, connect to Model Context Protocol (MCP) servers and return artefacts. A session can continue across turns, so an application can show progress and let a user steer the work. Agents API documentation.

The business value to test is reduced engineering effort around a task that genuinely needs several steps. Your team still defines what success means, which data is available and which actions require review. A managed runtime does not make an incomplete brief or unreliable source data disappear.

ILLUSTRATIVE WORKFLOW
AN INVESTIGATION WITH A REVIEWABLE RESULT
  1. Request
    scope + success criteria
  2. Session
    bounded working context
  3. Tools
    approved data + sandbox
  4. Verification
    tests + source checks
  5. Review
    accept or continue
Proposed application design. External writes pass through application permissions and approval rules.
WHERE TO START
01

An incident investigator

Start with a redacted alert, read-only logs and a runbook. Produce a timeline, likely causes and evidence links. A human owns the decision to change a live service.
02

A document reviewer

Compare a supplier submission with an approved requirements checklist. Return missing evidence and page references. Keep acceptance criteria outside the model and ask a reviewer to resolve disputed findings.
03

A business-data analyst

Answer a bounded question against approved warehouse views. Return the query, metric definitions and a chart. Reconcile totals to a trusted report before sharing the explanation.
WORKED EXAMPLE

Why did support demand rise?

Proposed pilot: an operations manager supplies a date range and asks why ticket volume increased. The agent reads permitted ticket summaries and product-release notes, groups the changes, and produces a source-linked explanation. Customer identifiers are removed unless needed for the approved question.

Verification runs outside the narrative: ticket counts must match the reporting system, categories must sum correctly, and every claimed release-related cause needs supporting evidence. Correlation is presented as a hypothesis until checked. Missing data produces an explicit limitation and a request for the missing source.

Evaluate time to a reviewed answer, arithmetic errors, unsupported explanations and total cost. Compare with a straightforward saved query and dashboard. The agent earns its place when the question varies enough that the fixed workflow becomes burdensome.

DEPLOYMENT AND ECONOMICS

Choose who runs each part.

OpenAI’s documentation separates the managed session from the execution environment. A self-hosted sandbox is not a self-hosted model or control plane. Map which prompts, tool results and artefacts cross each boundary before selecting the deployment. Runtime and environment overview.

Model tokens, tools and hosted containers are separate cost lines. Price a complete job, including waiting time, retries and review. Set time limits, stop abandoned environments and record spend by task. Published billing categories. Our AI observability guide explains the wider operational measurement.

Use a conventional worker for deterministic jobs. For cross-platform work owned by another agent, consider an explicit agent handoff. Multiple agents are useful when independent work can be verified and combined; adding them to a simple sequence introduces coordination overhead.

QUESTIONS
QUESTIONS — 3

The release changelog describes it as public beta as of 19 September 2026. Confirm current access, limits and contractual requirements before a production commitment.

Only through the tools and permissions your integration provides. Our proposed pilot starts with read-only investigation and introduces writes after explicit acceptance criteria and approval rules are tested.

No. Evaluate the full workflow against representative cases, including missing data, tool failures and a user changing the task midway through.

Keep reading

One workflow.
A measurable pilot.

Bring a repeated investigation or review task. We will define the data access, acceptance criteria and cost baseline before expanding autonomy.

CASE STUDIES

Shipped work.
Go and check it.

The work we can name, with the live site, our scope and the boundary made explicit. Select a project to see the evidence; each is a full case study, not a logo or a claim.

sonora.com
The Sonora homepage on desktop: a full-bleed dune landscape behind the headline “Transform Your Life with Sound”, with App Store and Google Play download buttons.
sonora.com — homepage, 1440×900 sonora.com →
Live Consumer wellness · Mobile + web

Sonora

Cognitive AI Ltd · 2026

A free sound-wellness app, described by its publisher as AI sound therapy that reads a short vocal sample at the start of a session and generates a soundscape for that moment. We designed and built the website and its backend, produced assets for the iOS and Android apps, and supported the application prototype.

Read the case study →

See every published project →

WHO WE HAVE BUILT FOR

Twenty-one years of applications, platforms and campaigns for names you know.

See all of our work →