$ curl llmtap.dev --same-model --10x-less

−90%

That's how much less you pay per token on the gpt-5.6 series versus official OpenAI pricing. The same frontier models (GPT, Claude, Gemini) behind a compatible API, billed per exact token. No subscription, no plan rate limits.

$3 llmtap credit = up to $30 at official GPT API pricing. Valid for 7 days after verification. No card.

terminal200 OK
$ curl https://www.llmtap.dev/api/v1/chat/completions \
    -H "Authorization: Bearer sk-live-your-key" \
    -H "Content-Type: application/json" \
    -d '{
      "model": "gpt-5.6-sol",
      "messages": [{"role": "user", "content": "hello"}]
    }'
gpt-5.6-sol $0.50/$3.00 gpt-5.6-terra $0.25/$1.50 gpt-5.5 $0.50/$3.00 gpt-5.4 $0.25/$1.50 codex-auto-review $0.25/$1.50 claude-opus-4.8 $0.80/$4.00 claude-sonnet-5 $0.32/$1.60 claude-haiku-4.5 $0.08/$0.32 claude-fable-5 $8.00/$40.00 gemini-3.1-pro $0.53/$3.20 gemini-3.5-flash $0.40/$2.40 USD per 1M tokens input/output

Verify the product before you trust it with traffic

Pricing, availability and compatibility are public. Inspect the operating surfaces directly before you send a meaningful workload.

Published catalog
11 routes with visible input and output pricing
Compatible APIs
Chat Completions, Responses and Anthropic Messages

Start in 4 steps

01

Register

Create your free account in 30 seconds. No card required.

02

Create API Key

Generate a named key from your dashboard. Copy it immediately — it shows only once.

03

Change Base URL

Point any OpenAI-compatible client to https://www.llmtap.dev/api/v1. One line, that is all.

04

Done

Use the same SDK, same code, same models — at a fraction of the cost.

How it works

01

Create your account and top up

Sign up with your email. Top up with USDT or USDC (BSC network) and your balance is credited automatically once the network confirms. Promotional redeem codes are also supported.

02

Generate your API key

From the dashboard you create named keys, revoke them anytime, and see per-request usage with exact tokens.

03

Change one line and call

The API is compatible with the OpenAI format: point base_url to llmtap, use your key, and your code stays the same.

Pricing per 1M tokens

Exact billing based on tokens reported by the upstream. Official is the list price of the corresponding provider; savings are computed on input.

modelproviderinputoutputofficial in/outsaving
gpt-5.6-solOpenAI$0.50$3.00$5.00/$30.0090%
gpt-5.6-terraOpenAI$0.25$1.50$2.50/$15.0090%
gpt-5.5OpenAI$0.50$3.00$5.00/$30.0090%
gpt-5.4OpenAI$0.25$1.50$2.50/$15.0090%
codex-auto-reviewOpenAI$0.25$1.50$2.50/$15.0090%
claude-opus-4.8Anthropic$0.80$4.00$5.00/$25.0084%
claude-sonnet-5Anthropic$0.32$1.60$2.00/$10.0084%
claude-haiku-4.5Anthropic$0.08$0.32$0.50/$2.5084%
claude-fable-5Anthropic$8.00$40.00$10.00/$50.0020%
gemini-3.1-proGoogle$0.53$3.20$2.00/$12.0073%
gemini-3.5-flashGoogle$0.40$2.40$1.50/$9.0073%

How much do you save?

Compare our resale pricing against official rates for your monthly usage.

$50$200$500$1,000$2,000
With llmtap
$200/mo
At official retail
$2,000/mo
You save
$1,800/mo(90%)

Based on a 10:1 input:output ratio — typical Cursor / Claude Code usage pattern.

Works with everything

Any OpenAI-compatible client. Change the base URL once, keep everything else.

CursorClaude CodeClineKilo CodeContinue.devOpenAI SDKAnthropic SDKLangChain
Universal config — paste this anywhere:
base_url = "https://www.llmtap.dev/v1"
api_key  = "sk-..."

You change one line. That's it.

If your code already talks to OpenAI, pointing it to llmtap means changing base_url and the key. Streaming, function calling and the error format behave the same.

And if the upstream fails, the charge is refunded automatically. You only pay for completed responses.

main.py
from openai import OpenAI

client = OpenAI(
    api_key="sk-live-your-key",
    base_url="https://www.llmtap.dev/api/v1",
)

res = client.chat.completions.create(
    model="claude-sonnet-5",
    messages=[{"role": "user", "content": "hello"}],
)
print(res.choices[0].message.content)

Honest questions

Why is it so cheap?

We buy capacity in bulk through subscription pools and route your requests to the official servers. Same model, shared cost. It is the same mechanism API proxies in Asia have used for years.

How do I know the real model answers, not a cheaper one?

The request goes to the official upstream unmodified. You can run your own benchmarks against the official API and compare: same weights, same results. Plus every request is logged with exact tokens in your dashboard.

What happens if the provider fails mid-request?

You don't pay. We lock an estimate before calling and settle against the real usage reported by the upstream. If the stream breaks or the provider errors, the hold is fully refunded to your balance automatically.

How do I pay?

With crypto: USDT or USDC on the BSC network, processed by NowPayments. Balance is credited only once the transaction confirms on-chain. No cards, no monthly subscription: you pay per token consumed.

Do I have to rewrite my code?

No. You change base_url and the API key. It works with the OpenAI SDK, with curl, and with any client that speaks the /v1/chat/completions format.

Are these the real models or knockoffs?

Real. Requests go to the official upstream unmodified. You can benchmark the same model against the official API and compare results: same weights, same output.

Why is it cheaper than official pricing?

We buy capacity in bulk through subscription pools and pass the savings to you. Same mechanism API proxies have used for years. No retail margin, no per-seat licensing overhead.

Can I use this in production?

Yes. We route through redundant upstream pools. If one provider degrades, requests fail over to the next available. Typical latency is within 5% of the official API. No request storage, metadata-only logging.

What about code privacy?

Requests are not stored or logged beyond metadata (model, tokens, timestamp) needed for billing. Prompt and response content is not retained. No training, no third-party sharing.

How do I connect to Cursor / Claude Code?

In Cursor: Settings → Models → Add OpenAI-compatible provider → base URL https://www.llmtap.dev/api/v1 + your key. In Claude Code: set ANTHROPIC_BASE_URL to https://www.llmtap.dev/api/v1 and use your key as the API token. Works with any OpenAI-compatible client.

Do you support subscriptions?

We are pay-as-you-go only for now: you pay per exact token consumed, nothing more. Subscriptions with monthly caps are on the roadmap.

$ ready-to-pay-less

Create your account, generate a key and call within the next 5 minutes.

create account