Doubleword
    Efficient Inference

    Frontier intelligence at a cost you control.

    Open-weight inference at up to 90% lower cost for long-horizon agents, batch pipelines, and high-volume tasks. Engineered for efficiency at every level of the stack.

    Frontier Quality, Open-Weight Economics

    Doubleword drives the 'Efficiency Frontier'

    Open-weight models, run Doubleword's most efficient inference stack give a better cost-performance trade-off than any alternative. Allowing our customers to take on more ambitious projects at the lowest possible price.

    Intelligence vs. cost

    Dotted rings are the cheapest option at each intelligence level.

    DoublewordOpenAIAnthropicEfficient frontier
    GPT-OSS-20B Qwen3.5-4B Qwen3.5-9B DeepSeek-V4-Flash-0731 GLM-5.3-Flash GLM-5.3 GPT-5.2 Claude Sonnet 4.6 (Max) GPT-5.6 Sol max Claude Opus 4.6 Claude Fable 5 GPT-6 Astra Claude Fable 5.1

    6 of the 10 models on the efficient frontier run on Doubleword. Hover any point for detail.

    See all models
    Model Target use case Intelligence Async / 1M out Batch / 1M out
    Kimi-K3 Agentic reasoning, long context High capability43.8intelligence score 43.8 $11.25 $7.50
    Qwen3.8-27B Coding and maths workflows High efficiency33.9intelligence score 33.9 $2.25 $1.50
    GLM-5.3-Flash Multimodal agentic coding High capability41.9intelligence score 41.9 $0.38 $0.25
    DeepSeek-V4-Flash-0731 High-throughput ETL, extraction High efficiency35intelligence score 35 $0.14 $0.09

    Intelligence via Artificial Analysis (Intelligence Index v4.3). Cost is one task of 1,000 tokens in / 500 out, at 24-hour batch rates for every provider.

    Testimonials

    Inference for token hungry teams

    Up to 90%
    lower cost than closed-source APIs
    4x
    Cheaper agentic inference than OpenRouter
    93%
    Average cache hit rates
    DetailnPlanOpenMedKenAIElement MaterialsUnaGo AIDataikuNXLUniversity of Colorado Boulder
    DetailnPlanOpenMedKenAIElement MaterialsUnaGo AIDataikuNXLUniversity of Colorado Boulder
    DetailnPlanOpenMedKenAIElement MaterialsUnaGo AIDataikuNXLUniversity of Colorado Boulder
    DetailnPlanOpenMedKenAIElement MaterialsUnaGo AIDataikuNXLUniversity of Colorado Boulder
    Total Cost Control

    Scale smarter with latency-based pricing.

    The same model at three speeds. Choose realtime, async or 24-hour batch, and pay less the longer you can wait.

    Realtime, async and batch inference pricing

    Model 44.9intelligence score 44.9

    Lowest latency, for interactive chat and live sessions.

    response = client.chat.completions.create(
    model="zai-org/GLM-5.3",
    messages=messages,
    service_tier="flex",
    )
    Cost per 1B in + 1B out
    $5.8K
    Full price, zero wait
    The closed source equivalent
    GPT-5.6 Sol max47.1intelligence score 47.1
    Costs $24K realtime.
    You save $18K
    Pay 4.1x less

    Switch in minutes.

    OpenAI Compatible API

    Switch instantly, with zero code rewrites required. Keep your existing OpenAI or Anthropic SDKs, update your base URL to api.doubleword.ai, and instantly run your workloads on our cost-aware open-weight models.

    Support from (human) assistants.

    Moving trillions of tokens? Our engineers give hands-on help with migration, prompt optimization and queue tuning.

    Prove the performance with evals.

    Don't guess on quality. Run our pre-built evaluation workflows to rigorously test output accuracy against your current production models before routing live traffic.

    INFERENCE ENGINEERING

    Pushing the frontier of Inference Engineering

    Doubleword contributes to the frontier of inference optimization research - building every level of the stack from kernels to scheduling to offer the most efficient inference for high volume workloads.

    Pricing that scales with your workload.

    Choose the model that best fits your business needs.

    Self-Service Inference

    Pay-as-you-go

    Whenever you need it, at the speed and cost that best fits your workload.

    Get started
    • Access to all major open-weight models
    • Usage-based pricing with prepaid credits
    • Prompt caching, tool calling and structured outputs
    • Realtime, async and batch inference
    • Zero Data Retention (ZDR) included
    • Community and standard support

    At-Scale Inference

    High-Throughput

    For specialized workloads requiring dedicated engineering and white-glove support.

    Talk to us
    • Everything available to self-service
    • Increased rate limits
    • Custom models
    • Workload-specific inference optimization (e.g., speculators) for improved throughput and reduced token costs
    • Region-locked data processing (e.g., US or EU-based inference)
    • Engineering support for workload optimization
    • Dedicated Slack channel with our engineering team
    • White-glove onboarding & Proof of Concept (PoC) support
    • Volume pricing
    • Signed MSA, DPA, and Custom SLAs

    FAQs

    Let there be tokens

    Try Doubleword's OpenAI compatible APIs and get frontier-level intelligence at a fraction of the cost.