The Infrastructure of Uncensored Intelligence

The uncensored AI API: Shannon models with no refusals and no filters, served from our own GPUs. OpenAI- and Anthropic-compatible — drop in the SDK you already use, one key, streaming, tool calling and reasoning on every endpoint.

21 uncensored models 256K context window from $0.50/1M 3 compatible APIs
shannon deepseekzaimoonshotainvidiaminimaxxiaomitencentpoolsidethinkingmachines

ماڈلز

SELF-HOSTED

Eight Shannon tiers and twelve hosted open-weight models, all behind the same endpoint. Each hosted id accepts what the model it imitates accepts — seven of the twelve are text-only. Model ids are stable — pin them in production.

Shannon 1.6 Lite

shannon-1.6-lite

روزمرہ کاموں کے لیے تیز، مؤثر جوابات

Context · 192K $3.90 / 1M

Shannon 1.6 Pro

shannon-1.6-pro

پیچیدہ مسائل کے لیے جدید ریزننگ

Context · 192K $7.80 / 1M

Shannon 2 Lite

shannon-2-lite

Shannon 2 Lite: fast think-draft-review pipeline with a visible reasoning trace. Everyday chat, summaries and Q&A.

Context · 192K $3.90 / 1M

Shannon 2 Pro

shannon-2-pro

Shannon 2 Pro: deeper think-draft-review passes for analysis, long-form writing and multi-step reasoning.

Context · 192K $5.85 / 1M

Shannon 3

shannon-3

Latest-generation Shannon model with fast reasoning at a low flat price.

Context · 192K $3.35 / 1M

Shannon 3 Pro

shannon-3-pro

Higher-capability Shannon 3 tier for complex, multi-step work.

Context · 192K $3.35 / 1M

Shannon 3.1

shannon-3.1

Shannon 3 on our own hardware, streaming at full speed with no rate gate.

Context · 192K $3.35 / 1M

Shannon 3.1 Pro

shannon-3.1-pro

Full-speed deep reasoning with a background research pass for hard questions.

Context · 192K $3.35 / 1M

Shannon Coder 1

shannon-coder-1

Code-and-tools specialist: agentic editing, long files, and strict tool loops.

Context · 128K $8.00 / 1M
3BIT-REAP

DeepSeek-V4-Pro-0813

DeepSeek-V4-Pro-0813-3BIT-REAP

Flagship-scale model for deep reasoning and long-form output.

Context · 256K Text in JSON schema
3BIT-REAP

GLM-5.2

GLM-5.2-3BIT-REAP

Strong generalist model with rich multilingual output.

Context · 256K Text in JSON schema
3BIT-REAP

Kimi-K3

Kimi-K3-3BIT-REAP

High-end model tuned for thorough, detailed responses.

Context · 256K Text + image in JSON schema
3BIT-REAP

Nemotron3Ultra

Nemotron3Ultra-3BIT-REAP

Efficient large model with strong reasoning per dollar.

Context · 256K Text in JSON schema
3BIT-REAP

MiniMax-M3

MiniMax-M3-3BIT-REAP

Fast, low-cost model for high-volume workloads.

Context · 256K Text + image in JSON schema
W4A16-AUTOROUND-REAP

DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP

Quick DeepSeek model balancing speed and quality.

Context · 256K Text in JSON schema
W4A16-AUTOROUND-REAP

Kimi-K2.6

Kimi-K2.6-W4A16-AUTOROUND-REAP

Capable model for everyday reasoning and drafting.

Context · 256K Text + image in JSON schema
W4A16-AUTOROUND-REAP

Laguna-S-2.1

Laguna-S-2.1-W4A16-AUTOROUND-REAP

Ultra-low-cost model for lightweight tasks.

Context · 256K Text in No structured output
W4A16-AUTOROUND-REAP

inkling

inkling-W4A16-AUTOROUND-REAP

Creative model with a distinctive writing voice.

Context · 256K Text + image in JSON object
W8A16

MiMo-V2.5-Pro

MiMo-V2.5-Pro-W8A16

Compact pro model for structured, precise output.

Context · 256K Text in JSON schema
W8A16

MiMo-V2.5

MiMo-V2.5-W8A16

Small, snappy model for simple tasks at minimal cost.

Context · 256K Text + image in JSON schema
W8A16

Hy3

Hy3-W8A16

Budget model for concise conversational replies.

Context · 256K Text in JSON schema

What each id accepts

The same values GET /v1/models returns. Every id streams and takes tool definitions; input and structured output vary, because a hosted id accepts exactly what the model it imitates accepts.

Model id Context Input Structured output Price / 1M
shannon-1.6-lite 192K Text + image in JSON schema $3.90 flat
shannon-1.6-pro 192K Text + image in JSON schema $7.80 flat
shannon-2-lite 192K Text + image in JSON schema $3.90 flat
shannon-2-pro 192K Text + image in JSON schema $5.85 flat
shannon-3 192K Text + image in JSON schema $3.35 flat
shannon-3-pro 192K Text + image in JSON schema $3.35 flat
shannon-3.1 192K Text + image in JSON schema $3.35 flat
shannon-3.1-pro 192K Text + image in JSON schema $3.35 flat
shannon-coder-1 128K Text in JSON schema $8.00 flat
DeepSeek-V4-Pro-0813-3BIT-REAP 256K Text in JSON schema $1.95 / $3.90
GLM-5.2-3BIT-REAP 256K Text in JSON schema $0.73 / $2.34
Kimi-K3-3BIT-REAP 256K Text + image in JSON schema $3.83 / $19.12
Nemotron3Ultra-3BIT-REAP 256K Text in JSON schema $0.75 / $3.30
MiniMax-M3-3BIT-REAP 256K Text + image in JSON schema $0.50 / $2.00
DeepSeek-V4-Flash-0731-W4A16-AUTOROUND-REAP 256K Text in JSON schema $0.50 / $2.00
Kimi-K2.6-W4A16-AUTOROUND-REAP 256K Text + image in JSON schema $0.78 / $3.67
Laguna-S-2.1-W4A16-AUTOROUND-REAP 256K Text in No structured output $0.50 / $2.00
inkling-W4A16-AUTOROUND-REAP 256K Text + image in JSON object $1.42 / $6.07
MiMo-V2.5-Pro-W8A16 256K Text in JSON schema $0.50 / $2.00
MiMo-V2.5-W8A16 256K Text + image in JSON schema $0.50 / $2.00
Hy3-W8A16 256K Text in JSON schema $0.50 / $2.00

Pricing

PER 1M TOKENS

Usage is billed against your token quota, valued at $5 per million quota tokens. Shannon models bill one flat rate for input and output; exact usage comes back on every response.

Model Price / 1M tokens
shannon-1.6-lite $3.90
shannon-1.6-pro $7.80
shannon-2-lite $3.90
shannon-2-pro $5.85
shannon-3 $3.35
shannon-3-pro $3.35
shannon-3.1 $3.35
shannon-3.1-pro $3.35
shannon-coder-1 $8.00

Hosted open-weight models

Model Variant Input / 1M Output / 1M
DeepSeek-V4-Pro-0813 3BIT-REAP $1.95 $3.90
GLM-5.2 3BIT-REAP $0.73 $2.34
Kimi-K3 3BIT-REAP $3.83 $19.12
Nemotron3Ultra 3BIT-REAP $0.75 $3.30
MiniMax-M3 3BIT-REAP $0.50 $2.00
DeepSeek-V4-Flash-0731 W4A16-AUTOROUND-REAP $0.50 $2.00
Kimi-K2.6 W4A16-AUTOROUND-REAP $0.78 $3.67
Laguna-S-2.1 W4A16-AUTOROUND-REAP $0.50 $2.00
inkling W4A16-AUTOROUND-REAP $1.42 $6.07
MiMo-V2.5-Pro W8A16 $0.50 $2.00
MiMo-V2.5 W8A16 $0.50 $2.00
Hy3 W8A16 $0.50 $2.00

Streaming responses include exact token usage in the final chunk. You are billed for the tokens you send and the reasoning and answer you receive — never for the pipeline's own rendering passes.

فوری آغاز

1 · Create a key 2 · Point your SDK at Shannon 3 · Ship
Python
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.shannon-ai.com/v1"
)

response = client.chat.completions.create(
    model="shannon-3",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Hello, Shannon!"}
    ],
    max_tokens=1024
)

print(response.choices[0].message.content)

جواب کا فارمیٹ

200 · JSON
{
  "id": "chatcmpl-abc123",
  "object": "chat.completion",
  "created": 1234567890,
  "model": "Shannon 1.6 Lite",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": "Hello! I'm Shannon, your AI assistant. How can I help you today?"
      },
      "finish_reason": "stop"
    }
  ],
  "usage": {
    "prompt_tokens": 25,
    "completion_tokens": 18,
    "total_tokens": 43
  }
}

API پلے گراؤنڈ

انٹرایکٹو

Try every model and endpoint in the browser with your own key — streaming output, request inspector, and generated code you can paste straight into your app.

Interactive API console

Chat with any model, three endpoint dialects, tool calls, live latency and cost.

Launch Playground

صلاحیتیں

Everything the chat product can do, exposed over the wire.

مطابق

ڈراپ‑اِن ریپلیسمنٹ

OpenAI اور Anthropic SDKs کے ساتھ کام کرتا ہے۔ بس base URL بدلیں۔

ٹولز

فنکشن کالنگ

ٹولز کی تعریف کریں، Shannon انہیں کال کرے گا۔ auto، forced، اور none موڈز سپورٹ۔

تلاش

بلٹ‑اِن ویب سرچ

ذرائع کے حوالوں کے ساتھ ریئل‑ٹائم ویب سرچ۔ خودکار دستیاب۔

JSON

اسٹرکچرڈ آؤٹ پٹس

قابلِ اعتماد ڈیٹا کے لیے JSON موڈ اور JSON Schema enforcement۔

ایجنٹک

ملٹی‑ٹرن ٹولز

خودکار فنکشن ایکزیکیوشن لوپس۔ ہر درخواست پر 10 iterations تک۔

تیز

اسٹریمنگ

ریئل‑ٹائم ٹوکن اسٹریمنگ کے لیے Server‑Sent Events۔

جائزہ

LIVE

Point your existing OpenAI or Anthropic SDK at Shannon and keep the same code. Every endpoint speaks the format you already use.

بنیادی URL

OpenAI‑مطابق

https://api.shannon-ai.com/v1/chat/completions

فنکشن کالنگ اور اسٹریمنگ کے ساتھ Chat Completions API استعمال کریں۔

بنیادی URL

Anthropic‑مطابق

https://api.shannon-ai.com/v1/messages

Claude Messages فارمیٹ ٹولز اور anthropic-version ہیڈر کے ساتھ۔

ہیڈرز

تصدیق

تصدیق: Bearer <aap-ki-key>

یا Claude‑اسٹائل کالز کے لیے X-API-Key کے ساتھ anthropic-version۔

رسائی

اسٹیٹس

عوامی ڈاکس - کال کرنے کے لیے کلید درکار ہے

اسٹریمنگ، فنکشن کالنگ، اسٹرکچرڈ آؤٹ پٹس، ویب سرچ۔

Before your first request

  • اپنے SDK کو Shannon کی طرف پوائنٹ کریں — baseURL کو اوپر دیے گئے OpenAI یا Anthropic endpoints پر سیٹ کریں۔
  • اپنی API کلید جوڑیں — OpenAI کالز کے لیے Bearer ٹوکن یا X-API-Key + anthropic-version استعمال کریں۔
  • ٹولز اور structured outputs فعال کریں — OpenAI tools/functions، JSON schema، اور built-in web_search سپورٹ کرتا ہے۔
  • استعمال ٹریک کریں — جب آپ سائن اِن ہوں تو اس صفحے پر ٹوکن اور سرچ کی کھپت دیکھیں۔

توثیق

One key works everywhere. OpenAI-style requests use a Bearer header; Anthropic-style requests use x-api-key.

OpenAI-compatible
Authorization: Bearer YOUR_API_KEY

Anthropic-compatible

Anthropic-compatible
x-api-key: YOUR_API_KEY
anthropic-version: 2023-06-01

فنکشن کالنگ

Python
from openai import OpenAI
import json

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.shannon-ai.com/v1"
)

# Define available tools/functions
tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get current weather for a location",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {
                        "type": "string",
                        "description": "City name, e.g., 'Tokyo'"
                    },
                    "unit": {
                        "type": "string",
                        "enum": ["celsius", "fahrenheit"]
                    }
                },
                "required": ["location"]
            }
        }
    }
]

response = client.chat.completions.create(
    model="shannon-3",
    messages=[{"role": "user", "content": "What's the weather in Tokyo?"}],
    tools=tools,
    tool_choice="auto"
)

# Check if model wants to call a function
if response.choices[0].message.tool_calls:
    tool_call = response.choices[0].message.tool_calls[0]
    print(f"Function: {tool_call.function.name}")
    print(f"Arguments: {tool_call.function.arguments}")

tool_choice

"auto" ماڈل طے کرتا ہے کہ فنکشن کال کرنی ہے یا نہیں (ڈیفالٹ)
"none" اس درخواست کے لیے فنکشن کالنگ بند کریں
{"type": "function", "function": {"name": "..."}} کسی مخصوص فنکشن کال کو مجبور کریں

فنکشن کال جواب

200 · JSON
{
  "id": "chatcmpl-xyz",
  "choices": [
    {
      "index": 0,
      "message": {
        "role": "assistant",
        "content": null,
        "tool_calls": [
          {
            "id": "call_abc123",
            "type": "function",
            "function": {
              "name": "get_weather",
              "arguments": "{\"location\": \"Tokyo\", \"unit\": \"celsius\"}"
            }
          }
        ]
      },
      "finish_reason": "tool_calls"
    }
  ]
}

ساختی آؤٹ پٹس

Python
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.shannon-ai.com/v1"
)

# Force JSON output with schema
response = client.chat.completions.create(
    model="shannon-3",
    messages=[
        {"role": "user", "content": "Extract: John Doe, 30 years old, engineer"}
    ],
    response_format={
        "type": "json_schema",
        "json_schema": {
            "name": "person_info",
            "schema": {
                "type": "object",
                "properties": {
                    "name": {"type": "string"},
                    "age": {"type": "integer"},
                    "occupation": {"type": "string"}
                },
                "required": ["name", "age", "occupation"]
            }
        }
    }
)

import json
data = json.loads(response.choices[0].message.content)
print(data)  # {"name": "John Doe", "age": 30, "occupation": "engineer"}

جواب فارمیٹ کے اختیارات

{"type": "json_object"} درست JSON آؤٹ پٹ لازمی کریں (بغیر مخصوص اسکیمہ)
{"type": "json_schema", "json_schema": {...}} آپ کے عین اسکیمہ سے ملتا آؤٹ پٹ لازمی کریں

اسٹریمنگ

Server-sent events, OpenAI chunk format. Thinking models stream reasoning_content deltas before the answer; the final chunk carries exact usage.

Python
from openai import OpenAI

client = OpenAI(
    api_key="YOUR_API_KEY",
    base_url="https://api.shannon-ai.com/v1"
)

# Enable streaming for real-time responses
# Thinking models stream reasoning_content first, then content
stream = client.chat.completions.create(
    model="GLM-5.2-3BIT-REAP",
    messages=[
        {"role": "user", "content": "Write a short poem about AI"}
    ],
    stream=True
)

for chunk in stream:
    delta = chunk.choices[0].delta
    if getattr(delta, "reasoning_content", None):
        print(delta.reasoning_content, end="", flush=True)  # thinking stream
    if delta.content:
        print(delta.content, end="", flush=True)

Responses API

NEW

POST /v1/responses — the OpenAI Responses dialect: instructions, input items, function_call / function_call_output for tool loops, reasoning summaries as output items. Same models, same key.

Python
from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")

response = client.responses.create(
    model="shannon-3",
    instructions="You are a concise assistant.",
    input="Summarize the three-way handshake in two sentences.",
    reasoning={"effort": "low"},
)
print(response.output_text)

# Tool loop: function_call items come back in response.output; answer them
# with function_call_output items on the next call.

Streaming emits response.created, reasoning_summary_text.delta, output_text.delta, function_call_arguments.delta and response.completed; a failed generation ends with response.failed. Non-streaming is the default.

Reasoning effort

NEW

Every thinking model streams its trace as reasoning_content deltas (or a thinking block) before the answer, and finish_reason: length tells you the answer hit max_tokens. The depth knob — reasoning_effort on /v1/chat/completions (off, low, medium, high), reasoning.effort on /v1/responses, thinking.budget_tokens on /v1/messages — is read by the hosted open-weight models; the Shannon tiers accept it and choose their own depth. GET /v1/models reports reasoning_effort per model.

Python
from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.shannon-ai.com/v1")

# reasoning_effort: off | low | medium | high. Read by the hosted open-weight
# models, where it sets how long the solve pass thinks. Shannon tiers accept the
# field and pick their own depth -- GET /v1/models reports which is which.
stream = client.chat.completions.create(
    model="GLM-5.2-3BIT-REAP",
    reasoning_effort="medium",
    stream=True,
    messages=[{"role": "user", "content": "Is 221 prime?"}],
)
for chunk in stream:
    delta = chunk.choices[0].delta
    if getattr(delta, "reasoning_content", None):
        print(delta.reasoning_content, end="", flush=True)   # thinking
    if delta.content:
        print(delta.content, end="", flush=True)             # answer

Anthropic

Drop-in for the Anthropic SDK — point it at our base URL and keep your Messages code.

https://api.shannon-ai.com/v1/messages
Python
import anthropic

client = anthropic.Anthropic(
    api_key="YOUR_API_KEY",
    base_url="https://api.shannon-ai.com"
)

response = client.messages.create(
    model="shannon-3",
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "Hello, Shannon!"}
    ],
    # Tool use (Anthropic format)
    tools=[{
        "name": "web_search",
        "description": "Search the web",
        "input_schema": {
            "type": "object",
            "properties": {
                "query": {"type": "string"}
            },
            "required": ["query"]
        }
    }]
)

print(response.content[0].text)

CLI coding tools

NEW

Use Shannon as the model behind Claude Code, Codex CLI and other agent CLIs.

Claude Code

Anthropic's official CLI coding agent. Point it at Shannon to use as your AI backend for reading, editing, and running code directly in your terminal.

ANTHROPIC_BASE_URL=https://api.shannon-ai.com ANTHROPIC_API_KEY=sk-YOUR_KEY claude

Codex CLI

OpenAI's open-source coding agent. Uses the Responses API for multi-turn tool use, file editing, and shell commands — all routed through Shannon.

OPENAI_BASE_URL=https://api.shannon-ai.com/v1 OPENAI_API_KEY=sk-YOUR_KEY codex

Claude Code

Shell
# Install Claude Code (requires Node.js 18+)
npm install -g @anthropic-ai/claude-code

# Connect to Shannon AI as backend
export ANTHROPIC_BASE_URL=https://api.shannon-ai.com
export ANTHROPIC_API_KEY=sk-YOUR_API_KEY

# Launch Claude Code in bare mode (no Anthropic account needed)
claude

# Or run a one-shot command
claude -p "Explain this codebase"

# Claude Code will use Shannon's Anthropic-compatible API
# for all AI operations: reading files, editing code,
# running tests, and multi-turn tool use.

Codex CLI

Shell
# Install Codex CLI
npm install -g @openai/codex

# Connect to Shannon AI as backend
export OPENAI_BASE_URL=https://api.shannon-ai.com/v1
export OPENAI_API_KEY=sk-YOUR_API_KEY

# Launch Codex
codex

# Or run a one-shot command
codex "fix the bug in main.py"

# Codex uses the Responses API (POST /v1/responses)
# Shannon handles tool calls including:
# - Reading and writing files
# - Running shell commands
# - Multi-turn function calling

SDKs

Any OpenAI or Anthropic SDK works out of the box.

Python

آفیشل OpenAI Python SDK - Shannon کے ساتھ کام کرتا ہے

pip install openai Documentation →

JavaScript / TypeScript

آفیشل OpenAI Node.js SDK - Shannon کے ساتھ کام کرتا ہے

npm install openai Documentation →

Go

OpenAI‑مطابق APIs کے لیے کمیونٹی Go کلائنٹ

go get github.com/sashabaranov/go-openai Documentation →

Ruby

OpenAI‑مطابق APIs کے لیے کمیونٹی Ruby کلائنٹ

gem install ruby-openai Documentation →

PHP

OpenAI‑مطابق APIs کے لیے کمیونٹی PHP کلائنٹ

composer require openai-php/client Documentation →

Rust

OpenAI‑مطابق APIs کے لیے Async Rust کلائنٹ

cargo add async-openai Documentation →

Python (Anthropic)

آفیشل Anthropic Python SDK - Shannon کے ساتھ کام کرتا ہے

pip install anthropic Documentation →

TypeScript (Anthropic)

آفیشل Anthropic TypeScript SDK - Shannon کے ساتھ کام کرتا ہے

npm install @anthropic-ai/sdk Documentation →

غلطیاں

Status Type Meaning
400 غلط درخواست غلط درخواست فارمیٹ یا پیرامیٹرز
401 غیر مجاز غلط یا غائب API کلید
429 کوٹہ ختم ٹوکن یا سرچ کوٹہ ختم
429 ریٹ لمٹ بہت زیادہ درخواستیں، رفتار کم کریں
500 سرور ایرر اندرونی خرابی، بعد میں دوبارہ کوشش کریں

Error body

4xx · JSON
{
  "error": {
    "message": "Invalid API key provided",
    "type": "authentication_error",
    "code": "invalid_api_key"
  }
}

تبدیلیوں کا لاگ

2.2.0

2026-03-28
  • نیا Claude Code support — use Shannon as your Anthropic backend for the official CLI coding agent
  • نیا Codex CLI support — full Responses API with multi-turn tool use for OpenAI's coding agent
  • بہتر بنایا Anthropic streaming format fixes — proper content_block lifecycle, tool_use deltas, toolu_ prefixes
  • بہتر بنایا Schema sanitization for Gemini — strips $schema, additionalProperties, $ref and other unsupported fields from tool schemas

2.1.0

2025-01-03
  • نیا Claude Code CLI انٹیگریشن کے لیے shannon-coder-1 ماڈل شامل کیا
  • نیا Coder ماڈل کے لیے کال‑بیسڈ کوٹہ سسٹم
  • بہتر بنایا فنکشن کالنگ کی قابلِ اعتمادیت بہتر کی

2.0.0

2024-12-15
  • نیا Anthropic Messages API مطابقت شامل کی
  • نیا ملٹی‑ٹرن ٹول ایکزیکیوشن (10 iterations تک)
  • نیا JSON Schema جواب فارمیٹ سپورٹ
  • بہتر بنایا بہتر حوالہ جات کے ساتھ ویب سرچ بہتر کی

1.5.0

2024-11-20
  • نیا پیچیدہ ریزننگ کے لیے shannon-deep-dapo ماڈل شامل کیا
  • نیا Built‑in web_search فنکشن
  • بہتر بنایا اسٹریمنگ جوابات کی لیٹنسی کم کی

1.0.0

2024-10-01
  • نیا ابتدائی API ریلیز
  • نیا OpenAI‑مطابق chat completions endpoint
  • نیا فنکشن کالنگ سپورٹ
  • نیا Server‑Sent Events کے ذریعے اسٹریمنگ

API کلید

GUEST

Sign in to view and manage your API key.

•••••••••••••••• Sign in

استعمال

Sign in to see your recent API activity.

API calls (30 days)
Billed tokens (30 days)

بنانے کے لیے تیار؟

Free tier included with every account. Your key is one click away.