Docs · quickstart
Quickstart.
Line Three serves POST /v1/chat/completions in the OpenAI chat completions format.
Pre-launch: api.linethree.ai does not answer requests until early access opens.
- Base URL
https://api.linethree.ai/v1- Endpoint
POST /v1/chat/completions- Auth
- Bearer or X-API-Key
Quickstart
Authentication
Every request carries your key in one of two headers. Use whichever your client supports.
Authorization: Bearer <key>, which is what OpenAI SDKs sendX-API-Key: <key>
Keep keys out of source control. Read them from an environment variable, as the examples below do.
First request
Use any model id from the list below. The examples use openai.gpt-oss-120b.
curl https://api.linethree.ai/v1/chat/completions \
-H "Authorization: Bearer $LINE_THREE_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai.gpt-oss-120b",
"messages": [
{"role": "user", "content": "Write table tests for parse_window()."}
]
}'
curl https://api.linethree.ai/v1/chat/completions \
-H "X-API-Key: $LINE_THREE_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "openai.gpt-oss-120b", "messages": [{"role": "user", "content": "…"}]}'
Models
Every model runs on Amazon Bedrock in US West (Oregon) (us-west-2), called through the Amazon Bedrock Runtime API (bedrock-runtime). Send Line Three’s model id exactly as shown; the gateway sends AWS’s bedrock-runtime id for that model, as listed on its AWS page (for gpt-oss-120b, openai.gpt-oss-120b-1:0; for the two Nemotron models, the same id as Line Three’s).
| Model | Model id | Context window | Input / output, per 1M tokens |
|---|---|---|---|
| gpt-oss-120b | openai.gpt-oss-120b | 128K tokens | $0.90 / $3.00 |
| NVIDIA Nemotron 3 Super 120B | nvidia.nemotron-super-3-120b | 256K tokens | $0.90 / $3.00 |
| NVIDIA Nemotron Nano 3 30B | nvidia.nemotron-nano-3-30b | 256K tokens | $0.25 / $0.90 |
What is supported
Line Three accepts the OpenAI chat completions format, not every OpenAI API. A client should work if it needs only what is listed here.
| Endpoint or feature | At launch | Details |
|---|---|---|
| POST /v1/chat/completions | Supported | model, messages, stream, temperature, top_p, max_tokens, stop, seed, response_format (JSON object or JSON schema) |
| Streaming | Supported | stream: true, with server-sent events; stream_options.include_usage for a final usage chunk |
| Tool calling | Partly | tools is passed to the model, and tool_choice is honoured only as "auto" (other values are dropped, as is parallel_tool_calls). The model’s calls come back in message.tool_calls; when streaming, each call arrives whole in one delta.tool_calls chunk rather than in argument fragments. |
| Coding agents | Not yet verified | No agent has been tested against Line Three yet. Agents that rely on streamed argument fragments or a forced tool_choice may not work. See Clients below. |
| GET /v1/models | Supported | lists the models your key can use |
| POST /v1/embeddings | Not offered | none of the launch models produces embeddings |
| POST /v1/completions (legacy) | Not supported | |
| Responses API, Assistants, files, batch, images, audio | Not supported |
Clients
Set the client’s base URL to https://api.linethree.ai/v1 and its API key to your Line Three key. With curl, the URL is the request itself, as in the first request above.
| Client | Kind | Status | Routes through vendor servers? |
|---|---|---|---|
| OpenAI Python SDK | Library | Compatible per the SDK’s documented base_url optionOpenAI(base_url="https://api.linethree.ai/v1", api_key=…) | No: it calls the base URL you set |
| OpenAI Node SDK | Library | Compatible per the SDK’s documented base_url optionnew OpenAI({ baseURL: 'https://api.linethree.ai/v1', apiKey: … }) | No: it calls the base URL you set |
| Cursor | Code editor | Not yet tested — tested results published before early access | Yes, per Cursor’s own docs (another website) |
| Continue | Editor extension | Not yet tested — tested results published before early access | Check the vendor |
| Cline | Editor extension | Not yet tested — tested results published before early access | Check the vendor |
| Open WebUI | Chat tool, run by your IT contractor | Not yet tested — tested results published before early access | Check the vendor |
| LibreChat | Chat tool, run by your IT contractor | Not yet tested — tested results published before early access | Check the vendor |
| AnythingLLM | Desktop chat tool | Not yet tested — tested results published before early access | Check the vendor |
Some editors relay your request through the editor maker’s own servers before it reaches the endpoint you set. If yours does, that company is in your request path too. Check your tool’s documentation. These tools keep their own chat history on the machine or server where they run; Line Three keeps no copy, but the tool may.
import os
from openai import OpenAI
client = OpenAI(
base_url="https://api.linethree.ai/v1",
api_key=os.environ["LINE_THREE_KEY"],
)
completion = client.chat.completions.create(
model="openai.gpt-oss-120b",
messages=[{"role": "user", "content": "Explain this module before I refactor it."}],
)
print(completion.choices[0].message.content)
import OpenAI from 'openai';
const client = new OpenAI({
baseURL: 'https://api.linethree.ai/v1',
apiKey: process.env.LINE_THREE_KEY,
});
const completion = await client.chat.completions.create({
model: 'openai.gpt-oss-120b',
messages: [{ role: 'user', content: 'Summarize this instruction into a checklist.' }],
});
console.log(completion.choices[0].message.content);
Streaming
Set "stream": true to receive the answer as it is written, as server-sent events ending in data: [DONE]. Add "stream_options": {"include_usage": true} for a final chunk with the token counts.
At the cap: a request that starts while the month’s cap has room is allowed to finish, even if it goes over the cap. The next request gets a 402. If a request is still running when the cap is reached or a prepaid balance reaches zero, Line Three absorbs that overrun and never bills it.
Errors
| Status | error.code | Meaning |
|---|---|---|
| 400 | invalid_request | The request body is not valid. The message says what is wrong. |
| 400 | context_length_exceeded | The request is longer than the model’s context window. |
| 401 | none | The key is missing or not accepted. This body isn’t the OpenAI shape: "error" is a string, "API key required" or "Invalid API key", with no code. |
| 402 | cap_reached | Prepaid balance exhausted, or the month’s cap, the key’s cap or the purchase order’s ceiling is reached. Do not retry; top up, raise the cap, or wait for the next month. |
| 429 | rate_limit_exceeded | Too many requests for this key this minute. Wait for the Retry-After seconds, then retry. |
| 503 | capacity, upstream_unavailable | Busy or briefly unavailable. Wait the number of seconds in the Retry-After header (5), then retry. |
| 504 | timeout | The request ran past the time limit. |
| 500 | internal_error | Something failed on our side. Retry with backoff. |
Errors on /v1/chat/completions have the OpenAI shape, with a request id to quote if you contact us:
{
"error": {
"message": "The service is at capacity. Please retry shortly.",
"type": "service_unavailable",
"code": "capacity",
"request_id": "1842"
}
}
Limits
- Per key
- Per-key caps are optional limits inside the account’s cap. The $50 minimum is per account, not per key.
- Rate
- Rate limits follow the cap, for card and purchase-order accounts alike: 60 requests a minute per key below a $1,000 monthly cap, and 600 a minute per key at $1,000 or more.
- Context
- gpt-oss-120b: 128K tokens; NVIDIA Nemotron 3 Super 120B: 256K tokens; NVIDIA Nemotron Nano 3 30B: 256K tokens.
- Request size
- Up to 10 MB of JSON per request.
- Timeout
- 10 minutes per request.