Docs · quickstart

Quickstart.

Line Three serves POST /v1/chat/completions in the OpenAI chat completions format.

Pre-launch: api.linethree.ai does not answer requests until early access opens.

Base URL
https://api.linethree.ai/v1
Endpoint
POST /v1/chat/completions
Auth
Bearer or X-API-Key

Quickstart

Authentication

Every request carries your key in one of two headers. Use whichever your client supports.

  • Authorization: Bearer <key>, which is what OpenAI SDKs send
  • X-API-Key: <key>

Keep keys out of source control. Read them from an environment variable, as the examples below do.

First request

Use any model id from the list below. The examples use openai.gpt-oss-120b.

curl · Bearer example
curl https://api.linethree.ai/v1/chat/completions \
  -H "Authorization: Bearer $LINE_THREE_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai.gpt-oss-120b",
    "messages": [
      {"role": "user", "content": "Write table tests for parse_window()."}
    ]
  }'
curl · X-API-Key example
curl https://api.linethree.ai/v1/chat/completions \
  -H "X-API-Key: $LINE_THREE_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "openai.gpt-oss-120b", "messages": [{"role": "user", "content": "…"}]}'

Models

Every model runs on Amazon Bedrock in US West (Oregon) (us-west-2), called through the Amazon Bedrock Runtime API (bedrock-runtime). Send Line Three’s model id exactly as shown; the gateway sends AWS’s bedrock-runtime id for that model, as listed on its AWS page (for gpt-oss-120b, openai.gpt-oss-120b-1:0; for the two Nemotron models, the same id as Line Three’s).

Models, their ids, context windows and rates
ModelModel idContext windowInput / output, per 1M tokens
gpt-oss-120bopenai.gpt-oss-120b128K tokens$0.90 / $3.00
NVIDIA Nemotron 3 Super 120Bnvidia.nemotron-super-3-120b256K tokens$0.90 / $3.00
NVIDIA Nemotron Nano 3 30Bnvidia.nemotron-nano-3-30b256K tokens$0.25 / $0.90

What is supported

Line Three accepts the OpenAI chat completions format, not every OpenAI API. A client should work if it needs only what is listed here.

Endpoints and features the API supports at launch
Endpoint or featureAt launchDetails
POST /v1/chat/completionsSupportedmodel, messages, stream, temperature, top_p, max_tokens, stop, seed, response_format (JSON object or JSON schema)
StreamingSupportedstream: true, with server-sent events; stream_options.include_usage for a final usage chunk
Tool callingPartlytools is passed to the model, and tool_choice is honoured only as "auto" (other values are dropped, as is parallel_tool_calls). The model’s calls come back in message.tool_calls; when streaming, each call arrives whole in one delta.tool_calls chunk rather than in argument fragments.
Coding agentsNot yet verifiedNo agent has been tested against Line Three yet. Agents that rely on streamed argument fragments or a forced tool_choice may not work. See Clients below.
GET /v1/modelsSupportedlists the models your key can use
POST /v1/embeddingsNot offerednone of the launch models produces embeddings
POST /v1/completions (legacy)Not supported
Responses API, Assistants, files, batch, images, audioNot supported

Clients

Set the client’s base URL to https://api.linethree.ai/v1 and its API key to your Line Three key. With curl, the URL is the request itself, as in the first request above.

Clients, what they are, and their test status
ClientKindStatusRoutes through vendor servers?
OpenAI Python SDKLibraryCompatible per the SDK’s documented base_url optionOpenAI(base_url="https://api.linethree.ai/v1", api_key=…)No: it calls the base URL you set
OpenAI Node SDKLibraryCompatible per the SDK’s documented base_url optionnew OpenAI({ baseURL: 'https://api.linethree.ai/v1', apiKey: … })No: it calls the base URL you set
CursorCode editorNot yet tested — tested results published before early accessYes, per Cursor’s own docs (another website)
ContinueEditor extensionNot yet tested — tested results published before early accessCheck the vendor
ClineEditor extensionNot yet tested — tested results published before early accessCheck the vendor
Open WebUIChat tool, run by your IT contractorNot yet tested — tested results published before early accessCheck the vendor
LibreChatChat tool, run by your IT contractorNot yet tested — tested results published before early accessCheck the vendor
AnythingLLMDesktop chat toolNot yet tested — tested results published before early accessCheck the vendor

Some editors relay your request through the editor maker’s own servers before it reaches the endpoint you set. If yours does, that company is in your request path too. Check your tool’s documentation. These tools keep their own chat history on the machine or server where they run; Line Three keeps no copy, but the tool may.

Python · openai example
import os
from openai import OpenAI

client = OpenAI(
    base_url="https://api.linethree.ai/v1",
    api_key=os.environ["LINE_THREE_KEY"],
)

completion = client.chat.completions.create(
    model="openai.gpt-oss-120b",
    messages=[{"role": "user", "content": "Explain this module before I refactor it."}],
)
print(completion.choices[0].message.content)
Node · openai example
import OpenAI from 'openai';

const client = new OpenAI({
  baseURL: 'https://api.linethree.ai/v1',
  apiKey: process.env.LINE_THREE_KEY,
});

const completion = await client.chat.completions.create({
  model: 'openai.gpt-oss-120b',
  messages: [{ role: 'user', content: 'Summarize this instruction into a checklist.' }],
});
console.log(completion.choices[0].message.content);

Streaming

Set "stream": true to receive the answer as it is written, as server-sent events ending in data: [DONE]. Add "stream_options": {"include_usage": true} for a final chunk with the token counts.

At the cap: a request that starts while the month’s cap has room is allowed to finish, even if it goes over the cap. The next request gets a 402. If a request is still running when the cap is reached or a prepaid balance reaches zero, Line Three absorbs that overrun and never bills it.

Errors

HTTP status codes Line Three returns, with the error code in the body
Statuserror.codeMeaning
400invalid_requestThe request body is not valid. The message says what is wrong.
400context_length_exceededThe request is longer than the model’s context window.
401noneThe key is missing or not accepted. This body isn’t the OpenAI shape: "error" is a string, "API key required" or "Invalid API key", with no code.
402cap_reachedPrepaid balance exhausted, or the month’s cap, the key’s cap or the purchase order’s ceiling is reached. Do not retry; top up, raise the cap, or wait for the next month.
429rate_limit_exceededToo many requests for this key this minute. Wait for the Retry-After seconds, then retry.
503capacity, upstream_unavailableBusy or briefly unavailable. Wait the number of seconds in the Retry-After header (5), then retry.
504timeoutThe request ran past the time limit.
500internal_errorSomething failed on our side. Retry with backoff.

Errors on /v1/chat/completions have the OpenAI shape, with a request id to quote if you contact us:

error body example
{
  "error": {
    "message": "The service is at capacity. Please retry shortly.",
    "type": "service_unavailable",
    "code": "capacity",
    "request_id": "1842"
  }
}

Limits

Per key
Per-key caps are optional limits inside the account’s cap. The $50 minimum is per account, not per key.
Rate
Rate limits follow the cap, for card and purchase-order accounts alike: 60 requests a minute per key below a $1,000 monthly cap, and 600 a minute per key at $1,000 or more.
Context
gpt-oss-120b: 128K tokens; NVIDIA Nemotron 3 Super 120B: 256K tokens; NVIDIA Nemotron Nano 3 30B: 256K tokens.
Request size
Up to 10 MB of JSON per request.
Timeout
10 minutes per request.

Opening to early-access teams first.