Documentation
Connect your apps to Somny
Create a key, choose a model, and send requests through OpenAI- or Anthropic-compatible endpoints.
Quick start
- Create a key. Sign in and open API keys in your dashboard. Copy the key when it appears; the full key is shown once.
- Add a custom connection. In your app, choose a custom OpenAI-compatible or Anthropic-compatible connection and enter the matching base URL below.
- Choose a model. Paste your Somny key and enter a model ID from the Models page. Use the exact ID.
https://somny.xyz/v1https://somny.xyzIf the app asks for a full endpoint, use /v1/chat/completions, /v1/responses, or /v1/messages on that host. Avoid adding /v1 twice. Keep the key in your app’s private API settings or your server environment, rather than public website code.
Make a request
Replace YOUR_SOMNY_KEY with your key and MODEL_ID with an available model. The SDK examples use the official OpenAI and Anthropic libraries; set SOMNY_API_KEY in your environment for the JavaScript versions.
List available models
curl https://somny.xyz/v1/models \
-H 'Authorization: Bearer YOUR_SOMNY_KEY'Chat Completions
from openai import OpenAI
client = OpenAI(
base_url="https://somny.xyz/v1",
api_key="YOUR_SOMNY_KEY",
)
reply = client.chat.completions.create(
model="MODEL_ID",
messages=[{"role": "user", "content": "Hello"}],
max_tokens=256,
)
print(reply.choices[0].message.content)Responses
The SDK examples reuse the OpenAI client from Chat Completions.
curl https://somny.xyz/v1/responses \
-H 'Authorization: Bearer YOUR_SOMNY_KEY' \
-H 'Content-Type: application/json' \
-d '{
"model": "MODEL_ID",
"input": "Hello",
"max_output_tokens": 256,
"store": false
}'Messages (Anthropic format)
curl https://somny.xyz/v1/messages \
-H 'x-api-key: YOUR_SOMNY_KEY' \
-H 'anthropic-version: 2023-06-01' \
-H 'Content-Type: application/json' \
-d '{
"model": "MODEL_ID",
"messages": [{"role": "user", "content": "Hello"}],
"max_tokens": 256
}'Streaming
Set "stream": true (or stream=True in Python) and add -N to curl. Read through the final event to receive final usage. An initial HTTP 200 does not guarantee that a stream completed.
Limits and features
Supported
- Text input and output
- Function tools, executed by your app
- Standard and streaming responses with final usage
- Chat Completions, Responses and Messages formats
Not available yet
- Image input
- Built-in hosted tools
- Stored-response retrieval and background requests
- Exact token counting
The Models page lists each model’s context and maximum requested output. The context includes input and output together and is checked using the provider’s token counting. Requests can be up to 16 MiB. A large context or output request still needs enough available allowance and can time out before reaching its maximum output.
Text and function-tool requests use stateless conversation history: include the relevant messages and tool results with each request. Support depends on the model and client.
Your allowance
Your plan has one daily allowance shared across your keys. It resets at midnight UTC; unused tokens expire. All paid plans share model access. The free plan has a smaller model selection.
The Models page and the Models screen in your dashboard list each model’s multipliers. Usage separates actual tokens from the weighted tokens deducted from your allowance:
Fresh input × input multiplier
+ cached input × cache-read multiplier
+ cache writes × cache-write multiplier
+ output × output multiplier
Long-context rate: GPT requests with more than 272,000 input tokens and Grok requests with more than 200,000 input tokens use 2× the usual weighted tokens for the entire request. The threshold counts fresh input, cache reads, and cache writes together. All input, cached tokens, and output are doubled; requests exactly at the threshold keep the usual rates.
Cache use comes from reported request usage, not an assumed cache percentage. Reasoning tokens are included in output and are not added a second time. Models with one flat multiplier use it for every category.
Reported input can include model instructions as well as your messages. Some models can exceed the requested output length. Your final allowance charge cannot exceed the amount reserved for that request.
Awaiting usage is a temporary hold. Somny reserves allowance before a request and replaces it with the measured charge when final usage is confirmed. A request can need more allowance up front than its eventual cached cost. Shorter context or a smaller output request can help when your balance is low.
If usage remains unconfirmed after 15 minutes, the hold is released against its original day and that request is not charged later. The request stays in your usage records. There are no automatic overage charges or token top-ups; an exhausted daily allowance resets the next day.
Somny keeps usage metadata, including request IDs, token counts, and timing. Somny does not save your prompts, responses, or function-tool payloads.
Errors and support
401Key not accepted- Check that you copied the full Somny key and that it has not been revoked.
400Request not supported- Check the error message for unsupported fields, model limits, or an invalid request format.
429Allowance or rate limit- Check your remaining allowance and pending requests. Follow a
Retry-Afterheader when one is provided. 503Temporarily unavailable- Check Status in your dashboard. Capacity may be full or the selected model may be paused.
If a connection drops or times out, check Usage before sending the request again. Work may have started even if you did not receive a complete answer. Keep the x-request-id response header when reporting a problem; it matches the request in your usage records.
For integrations, an optional Idempotency-Key prevents another dispatch with the same account-scoped key. A repeat returns 409 rather than replaying an answer, because Somny does not store generated responses.
Need help?
Join Somny’s Discord for support. Include your request ID when reporting an API problem, and keep API keys and payment details out of public channels.