AI Models. Sakana Fugu
Model orchestration / multi-agent / API

Sakana Fugu

multi-model orchestration through one API; Fugu and Fugu Ultra

Sakana Fugu, commercially released on June 22, 2026, is not a conventional standalone LLM. It presents one model endpoint while internally selecting, calling and coordinating a pool of model agents. The system decides whether to answer directly or split, verify and synthesize a more complex task. It is offered through an OpenAI-compatible API. Sakana positions Fugu as the lower-latency everyday option and Fugu Ultra as the quality-focused option for difficult, long-running work. The service is currently not offered to users in the EU or EEA.

Verified: 2026-06-23

Purchase decision (when to choose / when to avoid)

Choose if...

  • You operate outside the EU/EEA and want to submit complex work through one endpoint.
  • You do not want to design agent routing, roles, verification and synthesis yourself.
  • You accept a proprietary orchestration layer in exchange for simpler access to multiple models.

Avoid if...

  • Your users or organization are in the EU/EEA, where Sakana AI currently does not offer Fugu.
  • You must identify the model used for every request or audit the detailed routing trace.
  • You require independently reproduced benchmarks and a long production track record.

Cost in practice (scenarios)

Testing and light individual use

Standard at $20/month, described for occasional calls and small experiments.

  • service used outside the EU/EEA
  • a limited number of shorter sessions
Regular technical or research work

Pro at $100/month or Max at $200/month; published allowances are 10× and 20× Standard.

  • coding, review, research
  • longer weekly sessions
Fugu Ultra PAYG

$5 input / $30 output / $0.50 cached input per 1M tokens; above 272K context: $10 / $45 / $1.

  • fugu-ultra-20260615
  • agent fees are not stacked
These are estimates/scenarios (not an invoice). Actual cost depends on context length, number of users, limits and retention policies.

Deployment / data / enterprise

Deployment channels

  • OpenAI-compatible API through the Sakana AI console
  • Fugu - configurable pool with selected-agent opt-outs
  • Fugu Ultra - fixed full agent pool optimized for answer quality

Data policy

Training on data
Prompts and other usage data may be used to train and improve systems; users can opt out in the console. Opt-out is not retroactive.
Retention
Data is retained for as long as reasonably necessary; the policy does not provide one fixed period.
Data residency
Data may be transferred to locations including Japan and the United States. The service is not currently available in the EU/EEA.
Fugu supports selected-model opt-outs; Fugu Ultra uses its fixed full pool.

Enterprise readiness

Admin
Console, PAYG billing and per-request cost reporting.
SSO/SCIM
Not clearly specified on the public product page.
Audit
Token use and cost are reported per request; the underlying model-routing trace is not disclosed.
DPA
Confirm directly with Sakana AI before deployment.
Certifications
No clear certification list on the public product page.
For regulated use cases, the lack of model-level routing visibility may be material.

Best use cases

  • multi-step tasks that benefit from delegation, verification and synthesis
  • coding, code review, research analysis, and literature or patent investigations
  • teams that want a managed model pool behind one API instead of building their own agent router.

Strengths

  • One endpoint hides the operational complexity of model selection, delegation and result synthesis.
  • Standard Fugu lets customers exclude selected providers or models from its agent pool.
  • OpenAI API compatibility can reduce integration work compared with a custom multi-agent stack.

Weaknesses / risks

  • It is currently unavailable in the EU/EEA.
  • Sakana AI does not disclose which underlying models handled an individual request or the detailed routing trace.
  • The leading benchmark claims come from Sakana AI, while several comparison scores are provider-reported rather than independently reproduced.
  • Latency and cost can be less intuitive than calling one predetermined model directly.

Current models (examples)

  • Fugu - balances quality and latency and allows selected agents to be excluded from the pool.
  • Fugu Ultra (fugu-ultra-20260615) - quality-focused model using a fixed full agent pool.

Alternatives (if this model doesn't fit)