Set up your tool

Pick what you are using. Every one of them needs the same two things: the base URL https://llm.ccs.uky.edu/v1 and your key.

An empty reply is usually not an outage

Reasoning models spend part of your max_tokens budget thinking before they answer. If the budget runs out while the model is still thinking, you get back an empty reply, HTTP 200 with no error. The tell is completion_tokens coming back exactly equal to the max_tokens you asked for. There is no single safe number to set: try about 2,000 for a simple question and 8,000 or more for multi-step work, then raise it and retry if you see that signature. ccs/gpt-oss:120b is barely affected. On ccs/Qwen/Qwen3.6-35B-A3B-FP8 (thinking) the reply is never empty; instead the thinking text is cut off with no answer. Non-thinking and plain models are not affected. For multi-tool or agent work, prefer the thinking alias. The status page marks which models think first.

Which model should I use?

Several models are available, and any of them will answer an everyday request well: a question, a conversation, a page or two of text, a snippet of code. The choice below only really matters when your document is long, or when you are doing something specific like reading an image or building a search tool.

Everything in this section was checked with real requests on 2026-08-17. None of it is copied from a manufacturer's sheet.

The 30-second answer

If you want to…Use
Work with a long document or a large amount of codeQwen3.6 35B (no-think)
ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink
Build a tool-using assistant or agent (function/tool calls)Qwen3.6 35B (thinking)
ccs/Qwen/Qwen3.6-35B-A3B-FP8
Automate a plain, non-tool task (summarise, extract, transform)Qwen3.6 35B (no-think)
Ask a quick question and get a fast answerLlama 3.1 8B
ccs/llama3.1:8b
See careful step-by-step reasoning you can readQwen3.6 35B (thinking)
ccs/Qwen/Qwen3.6-35B-A3B-FP8
Ask about a picture, screenshot or chartQwen3-VL 8B
ccs/Qwen/Qwen3-VL-8B-Instruct-FP8
Build a search tool over your own documentsQwen3-Embedding 8B
ccs/Qwen/Qwen3-Embedding-8B

Qwen3.6 35B (no-think): "no-think" gives you just the answer, with no visible deliberation. It is not a weaker model; it is the same model with the step-by-step reasoning turned off so it replies faster, and it is the recommended default for chat, long documents and code. For a tool-calling agent, prefer the thinking variant (ccs/Qwen/Qwen3.6-35B-A3B-FP8): without the reasoning phase, the no-think alias can silently skip a tool call and answer from guesswork.

How much can I paste?

Every model has a limit on how much it can take in one request, and that limit counts your input only: every message in the conversation added together, including the system prompt and anything your program retrieved and pasted in. The length of the reply you ask for does not count towards it. We give the limit in pages of ordinary double-spaced text, because that is what you actually paste. These are rough equivalents, not exact conversions. The real figure depends on formatting and whether it is prose or code.

LimitRoughly the size of…If you are pasting code
about 25 pages (approximately 15,500 tokens)one journal article, a long email thread, a 10,000-word reportabout 1,000–1,500 lines
about 50 pages (approximately 33,000 tokens)two or three journal articles, a grant proposalabout 2,000–3,000 lines
about 100 pages (approximately 65,000 tokens)a Master's thesis, a short book, a full literature reviewabout 4,000–6,000 lines
about 200 pages (approximately 131,000 tokens)a PhD dissertation, a full manualabout 8,000–12,000 lines

The one thing everyone needs to know

Every model tells you when your text is too long. In almost every case you get a clear error naming the limit, nothing is sent, and you never get an answer built from a partial document. One narrow exception is noted below.

Two things follow from that. First, the limit counts your input only, not the reply you ask for. Lowering max_tokens will not bring a long document under the limit; only shortening the input will. Second, you no longer have to remember which models are “safe”: pick one whose limit fits your text, and if you go over you simply get an error instead of a wrong answer.

Rule of thumb: pick a model whose limit fits what you are pasting. If it is longer than about 25 pages, use Qwen3.6 35B (no-think). It holds about 200 pages (approximately 131,000 tokens), the most of anything here.

It adds up across a conversation. The limit applies to everything you send in one request, and a chat client resends the whole conversation every turn. System prompt, any text your program retrieved and pasted in, every earlier question and answer, and the new question are all counted together. This is why a chat that worked on turn 1 can start failing on turn 6 with nothing changed: the conversation itself grew past the limit. If that happens, start a new conversation or move to a model with a larger limit.

One narrow exception. For unusually dense input such as hashes, base64, or minified data, a prompt just under the limit can occasionally be trimmed with no error. If usage.prompt_tokens comes back at exactly the model’s limit, that happened and the front of your input was dropped. This matters most when the text you retrieve is itself dense (source code, logs, CSV, JSON): if you are building search over material like that, prefer a model with a much larger limit so you are never near the edge.

How much each model holds

These limits were re-measured with real requests on 2026-08-17. Every model now refuses over-length input with a clear error, so the only question is whether a model's limit is big enough for what you are pasting.

ModelYou can pasteIf you go over
Qwen3.6 35B (both versions)about 200 pages (approximately 131,000 tokens)Clear error. Nothing lost.
gpt-oss 120Babout 190 pages (approximately 124,000 tokens)Clear error. Nothing lost.
Qwen3-VL 8B (reads images too)about 100 pages (approximately 65,000 tokens)Clear error. Nothing lost.
Qwen3-Embedding 8Babout 50 pages per item (approximately 33,000 tokens)Clear error. Nothing lost.
DeepSeek-R1 32Babout 45 pages (approximately 28,000 tokens)Clear error. Nothing lost.
Qwen3 32Babout 25 pages (approximately 15,500 tokens)Clear error. Nothing lost.
Llama 3.1 8Babout 25 pages (approximately 15,500 tokens)Clear error. Nothing lost.

DeepSeek-R1 32B is a special case. Its limit of about 45 pages is not a hardware ceiling but a reliability one: past roughly that size it can no longer reliably find a fact in what you gave it, and may answer “not present” with full confidence. For a long document, use Qwen3.6 35B (either version) or gpt-oss 120B instead.

A note on each model

ModelGood forYou can pasteIf you go over
Qwen3.6 35B (no-think)
ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink
The default for chat, long documents and code. Very fast, and gives you the answer and nothing else. For tool-calling agents, use the thinking variant instead.about 200 pages (approximately 131,000 tokens)Clear error. Nothing lost.
Qwen3.6 35B (thinking)
ccs/Qwen/Qwen3.6-35B-A3B-FP8
The same model with its reasoning shown. Use it when you want to read how it reached the answer.about 200 pages (approximately 131,000 tokens)Clear error. Nothing lost.
Qwen3-VL 8B
ccs/Qwen/Qwen3-VL-8B-Instruct-FP8
The only model that can look at a picture: screenshots, charts, scanned pages, diagrams, photos.about 100 pages (approximately 65,000 tokens)Clear error. Nothing lost.
gpt-oss 120B
ccs/gpt-oss:120b
The largest model here. General chat and demanding work when quality matters more than speed.about 190 pages (approximately 124,000 tokens)Clear error. Nothing lost.
DeepSeek-R1 32B
ccs/deepseek-r1:32b
Problems that need working through step by step. Reliably handles about 45 pages; for a longer document use Qwen3.6 35B or gpt-oss 120B. Cannot do tool calling: sending tools returns an error, so it cannot be used by an agent.about 45 pages (approximately 28,000 tokens)Clear error. Nothing lost.
Qwen3 32B
ccs/qwen3:32b
Reasoning on shorter inputs.about 25 pages (approximately 15,500 tokens)Clear error. Nothing lost.
Llama 3.1 8B
ccs/llama3.1:8b
Quick questions, short summaries, rewriting and drafting. Usually replies in about a second.about 25 pages (approximately 15,500 tokens)Clear error. Nothing lost.
Qwen3-Embedding 8B
ccs/Qwen/Qwen3-Embedding-8B
Building search over your own documents. It turns text into numbers so you can search by meaning; it does not chat.about 50 pages per item (approximately 33,000 tokens)Clear error. Nothing lost.

A few quirks worth knowing

  • Llama 3.1 8B and search. It will happily accept a “turn this text into search numbers” request and hand back numbers that look real but are not. A search tool built on them quietly gives poor results, and nothing in the reply reveals the problem. Send search and “embedding” requests only to Qwen3-Embedding 8B.
  • Qwen3.6 35B (no-think) with automated tools. When you offer it several tools at once, it sometimes answers from memory instead of using one, and can state a made-up figure with full confidence. If your program relies on a tool to fetch live information, do not treat “no tool used” as “nothing was needed”; asking again and requiring the tool fixed it every time we tried.
  • Very short replies come back empty. The thinking version of Qwen3.6 35B and Qwen3 32B can return nothing at all if you ask for a very short reply, because they spend the whole budget thinking before they start. There is no single safe number: try about 2,000 tokens for a simple reply and 8,000 or more for multi-step work, then retry if it comes back empty. Or use a model that does not think first (though for multi-tool or agent work, prefer the thinking alias).
  • Qwen3.6 35B (thinking) is slower, because it reasons before it answers. As of 2026-08-28 it keeps that reasoning in a separate reasoning_content field and content holds only the answer; until that date the reasoning bled into the answer text and you could see a stray </think> in what came back. If you wrote code to strip that, it is no longer needed. The no-think version returns just the answer with no reasoning at all, which is another reason it is the default for scripts and automation.
  • Qwen3-VL 8B and Qwen3-Embedding 8B were the cleanest in testing: no problems found.

Writing code against the API? The service measures length in small units rather than pages: roughly 650 to a page, so the “about 25 pages” limit is about 15,500 of them. You never need to work this out; paste length is what matters.

Last tested 2026-08-17. Every limit and behaviour above was verified with real requests, not taken from documentation.

Overview

The CCS AI Inference Gateway provides access to large language models, hosted on CCS GPU hardware and, on a project key, commercial cloud models through Azure, all through one OpenAI-compatible API. This is the same API format used by OpenAI, and it is the de facto industry standard for LLM APIs.

Note: Access is granted per person on request. If you do not have an API key yet, ask for one at Locksmith. The service is in testing, so expect occasional interruptions.

How the pieces fit together

This host is a thin authenticating proxy. It hosts no models and performs no format translation. Your request is authenticated against your API key, then forwarded verbatim to the CCS inference gateway, which runs the local models on CCS GPU servers and relays project-key requests for azure/ models to the cloud. The response is streamed back to you unchanged.

Two consequences worth knowing:

  • API behaviour is defined upstream, not here. Which sampling parameters, tool-calling features, and response fields you get for a given model is decided by the inference gateway and the model itself. We do not add, drop, or rewrite parameters.
  • The model catalog changes. Models are added and removed as GPU hosts are re-provisioned. Always discover the current list with GET /v1/models rather than hard-coding names.

What does "OpenAI-compatible" mean for you?

It means any tool, library, or application that works with the OpenAI API works with our gateway. You just change two things:

  1. The base URL to https://llm.ccs.uky.edu/v1
  2. The API key to a key from Locksmith - your personal key, or a project key if you want that project's cloud models (see API keys)

No code changes, no special libraries, no adapter layers. The following tools and frameworks work out of the box:

OpenAI Python SDK OpenAI Node.js SDK OpenWebUI LangChain LlamaIndex Continue (VS Code) Cursor Jupyter AI Anthropic SDK / Claude Code LibreChat curl / HTTP clients TypeScript / Node.js Any OpenAI-compatible client

Base URL: https://llm.ccs.uky.edu/v1

Privacy: Your prompts and responses are never logged. We record metadata only - timestamp, model, status code, token counts, latency, client address, and the last four characters of the key; prompt and response content are never stored.

Getting Started

  1. Get a key from Locksmith - API keys are issued by CCS Key Management at locksmith.ccs.uky.edu. Sign in there with your institutional credentials via CILogon (University of Kentucky, ACCESS-CI, or another participating institution), open AI Inference Keys, and create your personal key. If you belong to a research project with a commercial-cloud budget, the same page also shows a project key for each such project. All keys issued by Locksmith begin with sk-.
  2. Copy your key - The full key is shown only once. Copy it immediately and store it securely.
  3. Point your client here - Set the base URL to https://llm.ccs.uky.edu/v1 and use your key as the API key. See the integration guides below.
  4. Discover the models - Call GET /v1/models (or check the Status Page) for the current catalog, then pass one of those ids as model.

Running an LLM from a cluster batch job (OOD "Launch an LLM") is different: those jobs are issued a short-lived ccs-llm- key automatically for the life of the SLURM job. You do not request those by hand. See API Key Management.

API Key Management

There are three kinds of key. All are sent the same way and work on every endpoint documented here; they differ in which models they can call and who pays. The key you configure in a tool decides both - there is no header, model-name prefix or setting to choose a project; the key is the choice.

Personal keyProject keyCluster job key
Issued by Locksmith, one per person Locksmith, one per project you belong to Automatically, by the OOD "Launch an LLM" job when it starts
Format sk-<random> sk-<random> ccs-llm-<random>
Models Local models (ccs/...) Local models and commercial cloud models through Azure (azure/...) Local models only
Who pays Nobody - local models are free Cloud calls draw on that project's dollar pool; local calls are free Nobody
Who it is for You: laptops, notebooks, editors, chat UIs You, when working for that project A single SLURM job on LCC / MCC / ECC
Lifetime Set by Locksmith policy when issued While you are a member of the project Bound to the job: revoked when the job ends, and swept by a reaper as a backstop
Revoke In Locksmith. Takes effect within about a minute; cannot be undone. In Locksmith, by you or the project PI. Takes effect within about a minute. Automatic at job exit; a CCS admin can also kill it immediately
Shown once The full key is displayed only at creation. Only a hash is stored, never the key itself.

This portal does not mint personal or project keys. Both come from Locksmith. Examples throughout this page use sk-YOUR_KEY_HERE as the placeholder for whichever sk- key you chose; if you are inside a cluster job, substitute the ccs-llm- key the job was given and everything else is identical.

Commercial cloud models through Azure cost money. Each project has a dollar pool that its PI arranged; every cloud call on a project key draws on it, and when the pool is spent, cloud calls on that key are refused until the PI adds funds (pools do not refill by themselves unless the PI set that up). Local models keep working on the same key throughout. The PI sees the balance in Locksmith. Cluster job keys can never reach the cloud models, by design: put your project key in the job if a batch job needs one.

Security: API keys are personal. Do not share them, commit them to git, or embed them in client-side code. Use environment variables or secret managers.

OpenAI API Compatibility

Our API implements the OpenAI chat completions specification. Any code written for the OpenAI API works by changing just the base URL and API key:

# Before: talking to OpenAI
client = OpenAI(api_key="<your OpenAI key>")

# After: talking to the CCS AI Inference Gateway - add base_url, swap the key
client = OpenAI(
    base_url="https://llm.ccs.uky.edu/v1",
    api_key="sk-YOUR_KEY"      # issued by locksmith.ccs.uky.edu
)

Everything else stays the same: client.chat.completions.create(), client.models.list(), streaming, etc. The only other change you will make is the model argument, which must name a model from our catalog (see Available Models) rather than an OpenAI model.

Note: CCS keys and OpenAI keys share the sk- prefix but are unrelated. A CCS key only works against https://llm.ccs.uky.edu.

API Reference

Base URL: https://llm.ccs.uky.edu/v1

Authentication: Authorization: Bearer sk-YOUR_KEY or x-api-key: sk-YOUR_KEY. Both headers are accepted on every endpoint, and any one key (personal or project) works for both the OpenAI-compatible (/v1/chat/completions) and Anthropic-compatible (/v1/messages) routes.

MethodEndpointDescriptionAuth
POST/v1/chat/completionsChat completions (primary endpoint)Required
POST/v1/messagesAnthropic-compatible Messages API (see Claude Code / Anthropic SDK)Required
POST/v1/completionsLegacy text completionsRequired
POST/v1/embeddingsGenerate embeddingsRequired
GET/v1/modelsList available modelsRequired
GET/health/livenessIs the service process upNone
GET/health/readinessIs the service up and the upstream gateway reachable (not per-model health; see the status page for that)None

Chat Completions Request Body

ParameterTypeRequiredDescription
modelstringYesModel ID (e.g., ccs/llama3.1:8b)
messagesarrayYesArray of {role, content} objects
temperaturefloatNoSampling temperature (0-2)
max_tokensintegerNoMaximum tokens to generate
top_pfloatNoNucleus sampling threshold
streambooleanNoEnable Server-Sent Events streaming
stoparrayNoStop sequences
tools, tool_choicearray / stringNoFunction calling. Not every model accepts these; see Available Models

This table lists the parameters in common use. The gateway does not filter the request body: any other OpenAI field you send is forwarded as-is to the inference gateway, which decides whether the chosen model honours it, ignores it, or rejects the request. See Known Limitations.

Response Format

{
    "id": "chatcmpl-1712345678000",
    "object": "chat.completion",
    "created": 1712345678,
    "model": "ccs/llama3.1:8b",
    "choices": [{
        "index": 0,
        "message": {"role": "assistant", "content": "..."},
        "finish_reason": "stop"
    }],
    "usage": {
        "prompt_tokens": 25,
        "completion_tokens": 150,
        "total_tokens": 175
    }
}

Embeddings

Embedding models convert text into numerical vectors for use in search, clustering, RAG (retrieval-augmented generation), and similarity tasks. The endpoint is POST /v1/embeddings, same format as OpenAI.

Pass encoding_format explicitly

Send "encoding_format": "float" on every request to ccs/Qwen/Qwen3-Embedding-8B. An earlier version of the service rejected requests that left it out; that has since been relaxed, but the OpenAI SDKs default to base64, so passing it explicitly is what makes your code do the same thing on every client and across service upgrades. It is always safe to send, and every example below does.

dimensions is ignored. Ask for 1,024 and you still get 4,096 back, with no error and no warning. If you need shorter vectors, truncate them yourself: these are Matryoshka embeddings, so the first N numbers of a vector are a valid shorter vector. Truncate every vector in an index to the same length, and re-normalise after truncating.

This model is embeddings only. Sending it to /v1/chat/completions will fail.

Request

curl -X POST https://llm.ccs.uky.edu/v1/embeddings \
  -H "Authorization: Bearer sk-YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ccs/Qwen/Qwen3-Embedding-8B",
    "input": "What is machine learning?",
    "encoding_format": "float"
  }'

The input field accepts either a single string or an array of strings. When you pass an array, the API returns one embedding per input in the same order, which is more efficient than making separate requests:

Keep a batch to 256 items or fewer. Very large float batches have been measured to take down a service worker, which returns errors to unrelated users, not just to you. When you are indexing a whole document collection, loop over batches of 256 rather than sending one enormous array.

curl -X POST https://llm.ccs.uky.edu/v1/embeddings \
  -H "Authorization: Bearer sk-YOUR_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ccs/Qwen/Qwen3-Embedding-8B",
    "input": ["First document", "Second document", "Third document"],
    "encoding_format": "float"
  }'

Response

{
    "object": "list",
    "data": [{
        "object": "embedding",
        "embedding": [0.0023, -0.0091, 0.0152, ...],
        "index": 0
    }],
    "model": "ccs/Qwen/Qwen3-Embedding-8B",
    "usage": {"prompt_tokens": 6, "total_tokens": 6}
}

Python Example

from openai import OpenAI

client = OpenAI(
    base_url="https://llm.ccs.uky.edu/v1",
    api_key="sk-YOUR_KEY"
)

# Generate embedding for a single string
response = client.embeddings.create(
    model="ccs/Qwen/Qwen3-Embedding-8B",
    input="What is machine learning?",
    encoding_format="float",      # always pass this explicitly
)

vector = response.data[0].embedding
print(f"Dimensions: {len(vector)}")
print(f"First 5: {vector[:5]}")

# Generate embeddings for multiple strings at once
response = client.embeddings.create(
    model="ccs/Qwen/Qwen3-Embedding-8B",
    input=["First document", "Second document", "Third document"],
    encoding_format="float",
)
for item in response.data:
    print(f"Index {item.index}: {len(item.embedding)} dimensions")

# Compare similarity of two texts
import numpy as np

def cosine_sim(a, b):
    return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))

def embed(text):
    r = client.embeddings.create(
        model="ccs/Qwen/Qwen3-Embedding-8B",
        input=text,
        encoding_format="float",
    )
    return r.data[0].embedding

cat, dog, plane = embed("cat"), embed("dog"), embed("airplane")

print(f"cat vs dog:      {cosine_sim(cat, dog):.4f}")
print(f"cat vs airplane: {cosine_sim(cat, plane):.4f}")
# Semantically related pairs score higher than unrelated ones.

Node.js Example

import OpenAI from "openai";

const client = new OpenAI({
    baseURL: "https://llm.ccs.uky.edu/v1",
    apiKey: "sk-YOUR_KEY",
});

// Single string
const single = await client.embeddings.create({
    model: "ccs/Qwen/Qwen3-Embedding-8B",
    input: "What is machine learning?",
    encoding_format: "float",   // always pass this explicitly
});
console.log(`Dimensions: ${single.data[0].embedding.length}`);

// Array of strings (batch)
const batch = await client.embeddings.create({
    model: "ccs/Qwen/Qwen3-Embedding-8B",
    input: ["First document", "Second document", "Third document"],
    encoding_format: "float",
});
for (const item of batch.data) {
    console.log(`Index ${item.index}: ${item.embedding.length} dimensions`);
}

Available Embedding Models

At the time of writing the catalog contains a single embedding model, ccs/Qwen/Qwen3-Embedding-8B. That can change: call GET /v1/models or check the Status Page, where embedding models are labelled "Embedding", for the current list.

Do not hard-code the vector length. Read it from the response (len(response.data[0].embedding)) so your index does not silently break if the hosted model changes.

Note: Embedding models use /v1/embeddings, not /v1/chat/completions. They return vectors, not text. Point an OpenAI-compatible embeddings client at /v1/embeddings to test embedding models.

Set the embedding model deliberately, then check it

A chat model asked for an embedding does not refuse. It returns a vector of a plausible length, correctly normalised, echoing back whatever model name you asked for. ccs/llama3.1:8b returns 4096 numbers, exactly like the real embedder does, so nothing in the response tells you the wrong model answered.

The failure that follows is quiet: a search index built on the wrong vectors still returns results, still looks fine on a handful of easy test documents, and simply retrieves the wrong document some of the time in production. A smoke test will not catch it.

So set the embedding model once, explicitly, in one place in your code, and assert on it. Send embedding requests only to ccs/Qwen/Qwen3-Embedding-8B. If you ever re-index, re-embed everything with the same model: vectors from two different models cannot be compared.

Budgeting context for RAG and web search

If you are building retrieval-augmented generation, or letting a model read web pages, the model's limit is the constraint that will bite you first, and it will bite in a confusing way. What you send each turn is not just the question:

system prompt + (chunks retrieved x tokens per chunk) + conversation so far + the question

All of it counts, and all of it is input, so lowering max_tokens does nothing. The conversation grows every turn, which is why a RAG chat frequently works perfectly on the first question and starts returning errors several questions later.

Working the arithmetic against the published limits, taking about 600 tokens of system prompt and question as overhead:

Retrieval settingsRoughly what you sendWhere it fits
4 chunks of 512 tokensabout 2,650Every model here.
8 chunks of 1,024 tokensabout 8,800Every model here.
8 chunks of 1,024, plus about five turns of conversationabout 12,800Fits the 25-page models on the first question, then fails as the conversation grows. Comfortable from DeepSeek-R1 32B upwards.
16 chunks of 1,024 tokensabout 17,000Over the limit on Qwen3 32B and Llama 3.1 8B. Fits DeepSeek-R1 32B and larger.
20 chunks of 1,500 tokensabout 30,600Needs Qwen3-VL 8B or larger; comfortable only on gpt-oss 120B and Qwen3.6 35B.

The simple answer: point retrieval and web-search traffic at ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink (about 200 pages, approximately 131,000 tokens) or ccs/gpt-oss:120b (about 190 pages, approximately 124,000 tokens). None of the settings above can overrun either one, and choosing one of them removes this whole class of failure rather than tuning around it. Reach for a smaller model for retrieval only if you have measured that you need its speed.

If you must stay on a smaller model, the levers are: retrieve fewer chunks, make the chunks shorter, or drop older turns from the conversation you resend. Those are the only three that work.

Do not treat these limits as permanent. They change as models are re-provisioned, so check the status page, which lists the current limit for every model and needs no key, before you settle on a chunk count for a long-lived index.

Settings for chat apps (Open WebUI and similar)

If you run your own chat application against this service, these are the settings to put in it. The defaults that ship with these apps are written for hosted commercial services with much larger limits, and leaving them alone is the usual cause of errors that look like the service is broken when it is not. A few of them can also degrade the service for other people, so they are worth setting deliberately.

SettingUse thisWhy
API base URLhttps://llm.ccs.uky.edu/v1The only address to use. Everything goes through here.
API keyYour own keyOne key identifies one person. Do not share it or put it in a page others can read.
Chat modelccs/Qwen/Qwen3.6-35B-A3B-FP8-nothinkFast, and its limit is the largest here, so long chats do not hit it.
Embedding modelccs/Qwen/Qwen3-Embedding-8BThe only real one. See the warning below; this is the setting people get wrong.
Embedding batch size256 or fewerLarger batches can take down a service worker and return errors to other users.
Chunk size500–1,000 tokensOrdinary RAG chunk sizes. The hard ceiling is 32,768 tokens per item.
Chunks retrieved (top-k)4–8 to startEvery retrieved chunk counts against the model's input limit. See the budget table.
Embedding dimensionsLeave unsetThe setting is ignored. You always get 4,096 back whatever you ask for.

The one setting that fails silently: the embedding model

If you type a chat model into the embedding field, you get no error. Measured on this service:

Model sent to /v1/embeddingsWhat comes back
ccs/Qwen/Qwen3-Embedding-8B4,096 numbers. Correct.
ccs/llama3.1:8b4,096 numbers, same shape, echoes the name you sent. Wrong, and undetectable.
ccs/deepseek-r1:32b5,120 numbers. Wrong.

The llama3.1:8b case is the dangerous one: it is the same length and the same shape as the real thing, so nothing in the reply tells you it is wrong. Your document search will build, run, and return plausible results while retrieving the wrong document a large fraction of the time. A quick test on a handful of documents will pass.

Set it once, deliberately, and check it. If you change it later, re-index everything: vectors made by two different models cannot be compared.

If your chat starts failing after a few messages

That is the input limit, not a fault. Your app resends the whole conversation every turn, so a chat that worked at the start can grow past the limit later, especially with documents attached. Start a new conversation, or use a model with a larger limit. Lowering the maximum reply length does not help, because the limit counts only what you send. See how much each model holds.

Python / OpenAI SDK

# pip install openai
from openai import OpenAI

client = OpenAI(
    base_url="https://llm.ccs.uky.edu/v1",
    api_key="sk-YOUR_KEY_HERE"
)

# Simple chat
response = client.chat.completions.create(
    model="ccs/llama3.1:8b",
    messages=[
        {"role": "system", "content": "You are a helpful research assistant."},
        {"role": "user", "content": "What is gradient descent?"}
    ],
    temperature=0.7,
    max_tokens=500
)
print(response.choices[0].message.content)
# Output: "Gradient descent is an optimization algorithm used to minimize
#          a function by iteratively moving in the direction of steepest
#          descent as defined by the negative of the gradient..."

print(f"Tokens used: {response.usage.total_tokens}")
# Output: Tokens used: 142

# List all available models
for model in client.models.list().data:
    print(model.id)
# Output:
#   ccs/deepseek-r1:32b
#   ccs/llama3.1:8b
#   ccs/qwen3:32b
#   ccs/Qwen/Qwen3-Embedding-8B
#   ...

Using Environment Variables (recommended)

# Set in your shell profile or .env file:
# export OPENAI_API_KEY="sk-YOUR_KEY_HERE"
# export OPENAI_BASE_URL="https://llm.ccs.uky.edu/v1"

# Then in Python - no credentials in code:
client = OpenAI()  # reads from environment automatically
response = client.chat.completions.create(
    model="ccs/llama3.1:8b",
    messages=[{"role": "user", "content": "Hello!"}]
)

TypeScript / Node.js

The official OpenAI Node.js / TypeScript SDK works with the CCS AI Inference Gateway. Install it with npm or your preferred package manager:

npm install openai

Chat Completion

import OpenAI from "openai";

const client = new OpenAI({
    baseURL: "https://llm.ccs.uky.edu/v1",
    apiKey: "sk-YOUR_KEY_HERE",
});

async function main() {
    const response = await client.chat.completions.create({
        model: "ccs/llama3.1:8b",
        messages: [
            { role: "system", content: "You are a helpful research assistant." },
            { role: "user", content: "What is gradient descent?" },
        ],
        temperature: 0.7,
        max_tokens: 500,
    });

    console.log(response.choices[0].message.content);
    console.log(`Tokens used: ${response.usage?.total_tokens}`);
}

main();

Streaming

import OpenAI from "openai";

const client = new OpenAI({
    baseURL: "https://llm.ccs.uky.edu/v1",
    apiKey: "sk-YOUR_KEY_HERE",
});

async function streamChat() {
    const stream = await client.chat.completions.create({
        model: "ccs/llama3.1:8b",
        messages: [{ role: "user", content: "Write a short essay on AI ethics." }],
        stream: true,
    });

    for await (const chunk of stream) {
        const content = chunk.choices[0]?.delta?.content;
        if (content) {
            process.stdout.write(content);
        }
    }
    console.log(); // newline at end
}

streamChat();

Embeddings

import OpenAI from "openai";

const client = new OpenAI({
    baseURL: "https://llm.ccs.uky.edu/v1",
    apiKey: "sk-YOUR_KEY_HERE",
});

async function embed() {
    const response = await client.embeddings.create({
        model: "ccs/Qwen/Qwen3-Embedding-8B",
        input: "What is machine learning?",
        encoding_format: "float",
    });

    const vector = response.data[0].embedding;
    console.log(`Dimensions: ${vector.length}`);
    console.log(`First 5: ${vector.slice(0, 5)}`);
}

embed();

TypeScript Type Hints

The SDK exports full type definitions. Use them for better autocompletion and type safety:

import OpenAI from "openai";
import type {
    ChatCompletion,
    ChatCompletionMessageParam,
    ChatCompletionCreateParamsNonStreaming,
} from "openai/resources/chat/completions";

const client = new OpenAI({
    baseURL: "https://llm.ccs.uky.edu/v1",
    apiKey: "sk-YOUR_KEY_HERE",
});

// Typed message array
const messages: ChatCompletionMessageParam[] = [
    { role: "system", content: "You are a helpful assistant." },
    { role: "user", content: "Explain TypeScript generics." },
];

// Typed request parameters
const params: ChatCompletionCreateParamsNonStreaming = {
    model: "ccs/llama3.1:8b",
    messages,
    temperature: 0.7,
    max_tokens: 500,
};

async function typedChat(): Promise<string | null> {
    const response: ChatCompletion = await client.chat.completions.create(params);
    return response.choices[0].message.content;
}

typedChat().then(console.log);

Using Environment Variables (recommended)

# Set in your shell profile or .env file:
# export OPENAI_API_KEY="sk-YOUR_KEY_HERE"
# export OPENAI_BASE_URL="https://llm.ccs.uky.edu/v1"

# Then in your code - no credentials needed:
const client = new OpenAI(); // reads from environment automatically

curl Examples

Chat Completion

curl -X POST https://llm.ccs.uky.edu/v1/chat/completions \
  -H "Authorization: Bearer sk-YOUR_KEY_HERE" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ccs/llama3.1:8b",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "What is the capital of Kentucky?"}
    ]
  }'

Response:

{
  "id": "chatcmpl-123",
  "object": "chat.completion",
  "created": 1712345678,
  "model": "ccs/llama3.1:8b",
  "choices": [{
    "index": 0,
    "message": {
      "role": "assistant",
      "content": "The capital of Kentucky is Frankfort."
    },
    "finish_reason": "stop"
  }],
  "usage": {
    "prompt_tokens": 28,
    "completion_tokens": 9,
    "total_tokens": 37
  }
}

Streaming

curl -N -X POST https://llm.ccs.uky.edu/v1/chat/completions \
  -H "Authorization: Bearer sk-YOUR_KEY_HERE" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ccs/llama3.1:8b",
    "messages": [{"role": "user", "content": "Write a haiku about computing."}],
    "stream": true
  }'

Response (Server-Sent Events):

data: {"id":"chatcmpl-456","object":"chat.completion.chunk","created":1712345678,"model":"ccs/llama3.1:8b","choices":[{"index":0,"delta":{"role":"assistant","content":"Silicon"},"finish_reason":null}]}

data: {"id":"chatcmpl-456","object":"chat.completion.chunk","created":1712345678,"model":"ccs/llama3.1:8b","choices":[{"index":0,"delta":{"content":" dreams"},"finish_reason":null}]}

data: {"id":"chatcmpl-456","object":"chat.completion.chunk","created":1712345678,"model":"ccs/llama3.1:8b","choices":[{"index":0,"delta":{"content":" awake"},"finish_reason":null}]}

...

data: {"id":"chatcmpl-456","object":"chat.completion.chunk","created":1712345678,"model":"ccs/llama3.1:8b","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":15,"completion_tokens":12,"total_tokens":27}}

data: [DONE]

List Models

curl https://llm.ccs.uky.edu/v1/models \
  -H "Authorization: Bearer sk-YOUR_KEY_HERE"

Response:

{
  "object": "list",
  "data": [
    {"id": "ccs/deepseek-r1:32b", "object": "model", "created": 1712345678, "owned_by": "ccs"},
    {"id": "ccs/gpt-oss:120b", "object": "model", "created": 1712345678, "owned_by": "ccs"},
    {"id": "ccs/llama3.1:8b", "object": "model", "created": 1712345678, "owned_by": "ccs"},
    {"id": "ccs/qwen3:32b", "object": "model", "created": 1712345678, "owned_by": "ccs"}
  ]
}

Abridged. This endpoint is the authoritative list of what you can pass as model, and it changes as GPU hosts are re-provisioned -- query it rather than copying ids out of this page.

Health Check (no key needed)

curl https://llm.ccs.uky.edu/health/readiness

Response:

{"checks": {"config": true, "db": true, "upstream": true}, "status": "ready", "upstream": {"db": "connected", "cache": null}}

Returns 200 with "status": "ready" when the inference gateway is reachable, 503 with "status": "not_ready" when it is not.

OpenWebUI

OpenWebUI is an open-source ChatGPT-like interface that supports custom OpenAI-compatible backends.

  1. Open your OpenWebUI instance
  2. Go to Admin Panel → Settings → Connections
  3. Under OpenAI API, click Add Connection
  4. Set URL to: https://llm.ccs.uky.edu/v1
  5. Set API Key to: your key from Locksmith (personal, or a project key to see that project's cloud models)
  6. Click the refresh icon to verify the connection
  7. Available models will appear in the model dropdown when starting a new chat

Tip: If you run OpenWebUI via Docker, you can also set OPENAI_API_BASE_URL and OPENAI_API_KEY as environment variables.

Jupyter AI (Jupyternaut Chat)

Jupyter AI adds a chat interface (Jupyternaut) and inline completions to JupyterLab. Each user configures their own settings, no shared environment variables needed.

Step-by-Step Setup

  1. Open the Jupyternaut chat panel in JupyterLab (left sidebar)
  2. Click the gear icon to open settings
  3. Fill in the fields exactly as shown below
  4. Scroll to the bottom and click Save Changes

Language Model

FieldExact Value
Completion modelOpenAI (general interface) :: * (select from dropdown)
Model IDccs/llama3.1:8b (no trailing spaces; any model from Status Page)
Base API URLhttps://llm.ccs.uky.edu/v1
OrganizationUKY (or leave blank)
Proxy(leave blank)

Embedding Model

FieldExact Value
Embedding modelOpenAI (general interface) :: * (select from dropdown)
Local model IDccs/Qwen/Qwen3-Embedding-8B (must be an embedding model, NOT a chat model)
Base API URLhttps://llm.ccs.uky.edu/v1

Inline Completions & API Keys

FieldExact Value
Inline completion modelNone (or configure same as language model for code autocomplete)
OPENAI_API_KEYYour key from Locksmith (e.g., sk-abc123...) - personal, or a project key for that project's cloud models. Inside an OOD cluster job, the job's ccs-llm-... key instead.

Common Errors

AssertionError: model_id was not specified
The Model ID field is empty. Type a model name like ccs/llama3.1:8b.
Model 'ccs/llama3.1:8b ' is not available (note the trailing space)
There is a space after the model name. Delete the trailing space.
Incorrect API key ... platform.openai.com
The request went to OpenAI instead of our gateway. The Base API URL is empty or was not saved. Set it to https://llm.ccs.uky.edu/v1 and click Save Changes again.
Embedding errors with ccs/llama3.1:8b
You used a chat model for the embedding model. Change it to ccs/Qwen/Qwen3-Embedding-8B.

Settings are stored per-user in ~/.local/share/jupyter/jupyter_ai/config.json. Each user on the cluster sets their own API key. Model ids are case-sensitive and contain slashes and colons; copy them exactly as GET /v1/models returns them.

Jupyter Notebooks (Python SDK)

# Install: pip install openai

from openai import OpenAI

client = OpenAI(
    base_url="https://llm.ccs.uky.edu/v1",
    api_key="sk-YOUR_KEY_HERE"
)

# --- List available models ---
models = client.models.list()
for m in models.data:
    print(m.id)

# --- Simple completion ---
response = client.chat.completions.create(
    model="ccs/llama3.1:8b",
    messages=[
        {"role": "system", "content": "You are a data science tutor."},
        {"role": "user", "content": "Explain PCA in simple terms."}
    ],
    temperature=0.7,
    max_tokens=500
)
print(response.choices[0].message.content)
print(f"Tokens used: {response.usage.total_tokens}")

# --- Streaming (good for long responses) ---
stream = client.chat.completions.create(
    model="ccs/llama3.1:8b",
    messages=[{"role": "user", "content": "Write a Python function to compute Fibonacci numbers."}],
    stream=True
)
for chunk in stream:
    # The stream ends with a usage-only chunk whose choices list is empty.
    content = chunk.choices[0].delta.content if chunk.choices else None
    if content:
        print(content, end="", flush=True)

# --- Multi-turn conversation ---
conversation = [
    {"role": "system", "content": "You are a helpful coding assistant."},
    {"role": "user", "content": "What is a decorator in Python?"},
]
r1 = client.chat.completions.create(model="ccs/llama3.1:8b", messages=conversation)
print(r1.choices[0].message.content)

conversation.append(r1.choices[0].message)
conversation.append({"role": "user", "content": "Can you give me a practical example?"})
r2 = client.chat.completions.create(model="ccs/llama3.1:8b", messages=conversation)
print(r2.choices[0].message.content)

VS Code / Continue Extension

Continue is an open-source AI code assistant for VS Code and JetBrains.

  1. Install the Continue extension from the VS Code marketplace
  2. Open the Continue config file: ~/.continue/config.json
  3. Add the CCS AI Inference Gateway as a model provider:
{
  "models": [
    {
      "title": "CCS AI Inference Gateway - Qwen3 32B",
      "provider": "openai",
      "model": "ccs/qwen3:32b",
      "apiBase": "https://llm.ccs.uky.edu/v1",
      "apiKey": "sk-YOUR_KEY_HERE"
    }
  ],
  "tabAutocompleteModel": {
    "title": "CCS Autocomplete",
    "provider": "openai",
    "model": "ccs/llama3.1:8b",
    "apiBase": "https://llm.ccs.uky.edu/v1",
    "apiKey": "sk-YOUR_KEY_HERE"
  }
}

Tip: You can add multiple models. Use a small, fast model for tab autocomplete and a larger one for chat. Check GET /v1/models or the Status Page for the current names. There is no dedicated code-completion model in the catalog today, so a general chat model is the right choice for both roles.

Cursor

Cursor is an AI-powered code editor that supports custom OpenAI-compatible endpoints.

  1. Open Cursor Settings (Cmd+, or Ctrl+,)
  2. Go to Models → OpenAI API Key
  3. Enter your Locksmith key
  4. Set Override OpenAI Base URL to: https://llm.ccs.uky.edu/v1
  5. Add a model name from our Status Page (e.g., ccs/llama3.1:8b)

Claude Code / Anthropic SDK

The gateway also speaks the Anthropic Messages API at https://llm.ccs.uky.edu/v1/messages. This lets Claude Code and the Anthropic Python/TypeScript SDKs target the gateway directly with the same key you use for the OpenAI endpoint. Pick whichever wire format your client sends natively. This route is served by the upstream inference gateway, not translated here.

Current scope

Text and vision prompts, streaming and non-streaming, and tool use (single and parallel, including tools exposed to the model over MCP) are all supported. How well tool use works depends entirely on the model you choose, not on this portal. ccs/qwen3:32b and the thinking ccs/Qwen/Qwen3.6-35B-A3B-FP8 variant are reasonable starting points for agentic work. Prefer the thinking variant over the -nothink alias for multi-tool loops: without the reasoning phase, -nothink can silently skip a tool call and fabricate the result. ccs/deepseek-r1:32b does not support tool calling and returns 400 if you send tools.

Do not send tool_choice: "none" to either ccs/Qwen/Qwen3.6-35B-A3B-FP8 alias. It is not honoured: instead of suppressing tool use, the model's raw tool markup is returned to you inside content with finish_reason: "stop", which your parser will read as an ordinary answer. If you do not want a tool called on a given turn, leave tools out of that request entirely. The opposite direction works correctly and is the recommended fix for a skipped call: tool_choice: "required" on turns that must hit a tool.

Anthropic's signed thinking blocks cannot be reproduced by non-Claude models, so that request field does not behave as it would against Anthropic. Reasoning models such as ccs/deepseek-r1:32b emit their reasoning inline or in a reasoning field instead, depending on the model and the endpoint.

Claude Code

Set two environment variables before launching claude:

export ANTHROPIC_BASE_URL=https://llm.ccs.uky.edu
export ANTHROPIC_API_KEY=sk-YOUR_KEY_HERE

claude

Claude Code will call ${ANTHROPIC_BASE_URL}/v1/messages and send your key as x-api-key. Streaming and tool use are both supported, so the fully agentic flow (Read, Write, Bash, etc.) works. Pick a tool-capable model (e.g., ccs/qwen3:32b or ccs/Qwen/Qwen3.6-35B-A3B-FP8, the thinking variant, which is the safer choice for tool-heavy agent loops). Do not use ccs/deepseek-r1:32b here: it rejects requests that carry tools.

Anthropic Python SDK

# pip install anthropic
from anthropic import Anthropic

client = Anthropic(
    base_url="https://llm.ccs.uky.edu",
    api_key="sk-YOUR_KEY_HERE",
)

msg = client.messages.create(
    model="ccs/llama3.1:8b",
    max_tokens=512,
    system="You are a helpful research assistant.",
    messages=[
        {"role": "user", "content": "What is gradient descent?"}
    ],
)
print(msg.content[0].text)
print(f"Input tokens: {msg.usage.input_tokens}, Output tokens: {msg.usage.output_tokens}")

curl

curl -X POST https://llm.ccs.uky.edu/v1/messages \
  -H "x-api-key: sk-YOUR_KEY_HERE" \
  -H "anthropic-version: 2023-06-01" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ccs/llama3.1:8b",
    "max_tokens": 256,
    "messages": [
      {"role": "user", "content": "What is the capital of Kentucky?"}
    ]
  }'

Key differences from OpenAI /v1/chat/completions

AspectOpenAI routeAnthropic route
Endpoint/v1/chat/completions/v1/messages
Auth headerAuthorization: Bearer ...x-api-key: ... (Bearer also accepted)
System promptMessage with role:"system"Top-level system field
max_tokensOptionalRequired
Stop sequencesstopstop_sequences
Usageprompt_tokens / completion_tokensinput_tokens / output_tokens
Imagesimage_url data URLimage block with source.type:"base64"

The model field must match an id returned by GET /v1/models (also shown on the Status Page). Anthropic model names like claude-opus-4-7 are not remapped. Send the CCS model id (e.g., ccs/llama3.1:8b).

LangChain

# pip install langchain-openai

from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    base_url="https://llm.ccs.uky.edu/v1",
    api_key="sk-YOUR_KEY_HERE",
    model="ccs/llama3.1:8b",
    temperature=0.7,
    max_tokens=500
)

# Simple invocation
response = llm.invoke("What is retrieval-augmented generation?")
print(response.content)

# With message history
from langchain_core.messages import HumanMessage, SystemMessage

messages = [
    SystemMessage(content="You are an expert on natural language processing."),
    HumanMessage(content="Compare BERT and GPT architectures.")
]
response = llm.invoke(messages)
print(response.content)

# Streaming
for chunk in llm.stream("Explain attention mechanisms step by step"):
    print(chunk.content, end="", flush=True)

LlamaIndex

# pip install llama-index-llms-openai-like

from llama_index.llms.openai_like import OpenAILike

llm = OpenAILike(
    api_base="https://llm.ccs.uky.edu/v1",
    api_key="sk-YOUR_KEY_HERE",
    model="ccs/llama3.1:8b",
    is_chat_model=True,
    temperature=0.7,
    max_tokens=500
)

# Simple completion
response = llm.complete("Explain vector databases in one paragraph.")
print(response.text)

# Chat
from llama_index.core.llms import ChatMessage

messages = [
    ChatMessage(role="system", content="You are a helpful assistant."),
    ChatMessage(role="user", content="What are embeddings?"),
]
response = llm.chat(messages)
print(response.message.content)

Streaming Responses

Streaming delivers tokens as they are generated, providing a much better user experience for long responses. Set "stream": true in your request.

Python (OpenAI SDK)

stream = client.chat.completions.create(
    model="ccs/llama3.1:8b",
    messages=[{"role": "user", "content": "Write a short essay on AI ethics."}],
    stream=True
)

for chunk in stream:
    # The stream ends with a usage-only chunk whose choices list is empty.
    content = chunk.choices[0].delta.content if chunk.choices else None
    if content:
        print(content, end="", flush=True)
print()  # newline at end

Node.js (OpenAI SDK)

const stream = await client.chat.completions.create({
    model: "ccs/llama3.1:8b",
    messages: [{ role: "user", content: "Write a short essay on AI ethics." }],
    stream: true,
});

for await (const chunk of stream) {
    const content = chunk.choices[0]?.delta?.content;
    if (content) {
        process.stdout.write(content);
    }
}
console.log(); // newline at end

Node.js (fetch)

const response = await fetch('https://llm.ccs.uky.edu/v1/chat/completions', {
    method: 'POST',
    headers: {
        'Authorization': 'Bearer sk-YOUR_KEY',
        'Content-Type': 'application/json'
    },
    body: JSON.stringify({
        model: 'ccs/llama3.1:8b',
        messages: [{role: 'user', content: 'Hello!'}],
        stream: true
    })
});

const reader = response.body.getReader();
const decoder = new TextDecoder();

while (true) {
    const {done, value} = await reader.read();
    if (done) break;
    const chunk = decoder.decode(value);
    // Parse SSE lines: "data: {...}\n\n"
    for (const line of chunk.split('\n')) {
        if (line.startsWith('data: ') && line !== 'data: [DONE]') {
            const data = JSON.parse(line.slice(6));
            process.stdout.write(data.choices[0]?.delta?.content || '');
        }
    }
}

Vision Models

Vision models accept images alongside text in the messages array. Images are sent as base64-encoded data URLs using the standard OpenAI image_url content-part format. The vision model in the catalog at the time of writing is ccs/Qwen/Qwen3-VL-8B-Instruct-FP8; models labelled "Vision" on the Status Page are the current set.

Message Format for Images

Instead of a plain string, set the content field to an array containing text and image parts:

{
    "model": "ccs/Qwen/Qwen3-VL-8B-Instruct-FP8",
    "messages": [
        {
            "role": "user",
            "content": [
                {
                    "type": "text",
                    "text": "What is in this image?"
                },
                {
                    "type": "image_url",
                    "image_url": {
                        "url": "data:image/png;base64,iVBORw0KGgo..."
                    }
                }
            ]
        }
    ]
}

Python Example

import base64
from openai import OpenAI

client = OpenAI(
    base_url="https://llm.ccs.uky.edu/v1",
    api_key="sk-YOUR_KEY_HERE"
)

# Read and encode a local image
with open("photo.png", "rb") as f:
    image_b64 = base64.b64encode(f.read()).decode("utf-8")

response = client.chat.completions.create(
    model="ccs/Qwen/Qwen3-VL-8B-Instruct-FP8",
    messages=[
        {
            "role": "user",
            "content": [
                {"type": "text", "text": "Describe this image in detail."},
                {
                    "type": "image_url",
                    "image_url": {
                        "url": f"data:image/png;base64,{image_b64}"
                    },
                },
            ],
        }
    ],
    max_tokens=500,
)
print(response.choices[0].message.content)
# Output: "The image shows a bar chart with three colored bars representing
#          quarterly revenue data. The x-axis labels show Q1, Q2, and Q3..."

Node.js Example

import OpenAI from "openai";
import { readFileSync } from "node:fs";

const client = new OpenAI({
    baseURL: "https://llm.ccs.uky.edu/v1",
    apiKey: "sk-YOUR_KEY_HERE",
});

// Read and encode a local image
const imageBuffer = readFileSync("photo.png");
const imageB64 = imageBuffer.toString("base64");

const response = await client.chat.completions.create({
    model: "ccs/Qwen/Qwen3-VL-8B-Instruct-FP8",
    messages: [
        {
            role: "user",
            content: [
                { type: "text", text: "Describe this image in detail." },
                {
                    type: "image_url",
                    image_url: {
                        url: `data:image/png;base64,${imageB64}`,
                    },
                },
            ],
        },
    ],
    max_tokens: 500,
});

console.log(response.choices[0].message.content);
// Output: "The image shows a bar chart with three colored bars..."

curl Example

# Encode the image (macOS / Linux)
IMAGE_B64=$(base64 photo.png | tr -d '\n')

curl -X POST https://llm.ccs.uky.edu/v1/chat/completions \
  -H "Authorization: Bearer sk-YOUR_KEY_HERE" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ccs/Qwen/Qwen3-VL-8B-Instruct-FP8",
    "messages": [{
      "role": "user",
      "content": [
        {"type": "text", "text": "What is in this image?"},
        {"type": "image_url", "image_url": {"url": "data:image/png;base64,'"$IMAGE_B64"'"}}
      ]
    }]
  }'

Note: Only models with vision capabilities can process images. Sending images to a text-only model will result in an error. Check the Status Page to see which models support vision.

Model Types

The catalog contains several categories of model, each suited for different tasks. Understanding the differences helps you choose the right one. The examples below are illustrative of the catalog as it stands today; it changes, so confirm with GET /v1/models.

Chat Models

General-purpose language models designed for text generation, conversation, question answering, summarization, and reasoning tasks. These are the most commonly used models and are a good default choice.

Examples: ccs/llama3.1:8b, ccs/qwen3:32b, ccs/gpt-oss:120b

Vision Models

Multimodal models that can accept both text and images as input. Use these when you need a model to analyze, describe, or answer questions about images. Images are sent as base64-encoded data in the messages array (see the Vision Models section above for the exact format).

Example: ccs/Qwen/Qwen3-VL-8B-Instruct-FP8

Reasoning Models

Models that produce an explicit chain of thought before their final answer. This generally improves accuracy on multi-step problems at the cost of latency and tokens. Depending on the model and endpoint, the reasoning arrives either inline in the text or in a separate reasoning field (delta.reasoning when streaming); clients such as Open WebUI render it in a collapsible section.

Example: ccs/deepseek-r1:32b. Note it does not support tool calling: sending tools returns 400.

Some models ship a paired -nothink variant, for example ccs/Qwen/Qwen3.6-35B-A3B-FP8 and ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink. These are the same weights; the -nothink id simply has chat_template_kwargs.enable_thinking = false applied, suppressing the thinking phase. Prefer -nothink for latency-sensitive and non-tool calls, where the thinking output is just noise. For agent and tool-calling loops (including MCP-driven agents) prefer the thinking variant instead: in testing, -nothink skipped a tool call and fabricated the answer in about one in five multi-tool turns (the thinking variant: none).

Code Work

There is no dedicated code-completion model in the catalog today. General chat models handle code generation, completion, and debugging, and are what the Continue and Cursor guides above configure. If you need a code-specialised model, ask CCS.

Embedding Models

These models convert text into fixed-length numerical vectors (embeddings). They do not generate text. Instead, the vectors they produce capture the semantic meaning of the input, so similar texts have vectors that are close together in the embedding space. Embeddings are used for:

  • Semantic search: find documents similar to a query
  • RAG (retrieval-augmented generation): retrieve relevant context before generating an answer
  • Clustering and classification: group or categorize text by meaning

Embedding models use the /v1/embeddings endpoint, not /v1/chat/completions. See the Embeddings section for usage details.

Example: ccs/Qwen/Qwen3-Embedding-8B, which requires encoding_format: "float" on every request.

Send embeddings only to the embedding model. Some chat models will answer /v1/embeddings with a 200 and a correctly-shaped vector that is not a usable embedding. For example, ccs/llama3.1:8b returns a 4096-dimension vector, the same size as the real embedder, so a client cannot tell by shape. Always use ccs/Qwen/Qwen3-Embedding-8B for embeddings.

Context Window Sizes

How much text each model can take in one request differs from model to model, and is covered plainly in Which model should I use?, with a figure in pages for each one. Every model now refuses over-length input with a clear error naming the limit (nothing is sent and nothing is silently dropped) so the only thing to check before you paste is that a model's limit is big enough for your text.

Available Models

Local models (ccs/) run on CCS GPU servers behind the inference gateway; azure/ models are commercial cloud models reachable on a project key. The catalog is not fixed: models are added and removed as GPU hosts are re-provisioned, so treat any list printed in documentation (including this page) as a snapshot.

The authoritative list is GET /v1/models. Query it at runtime rather than hard-coding ids, and surface the result in your app's model picker where you can. The Status Page shows the same list with live up/down state.

Model ids are case-sensitive and contain slashes and colons (ccs/Qwen/Qwen3-Embedding-8B). Copy them exactly, with no trailing whitespace.

Model-specific behaviour worth knowing

ModelWhat to know
ccs/deepseek-r1:32b Reasoning model. Does not support tool calling: a request carrying tools returns 400. Use a different model for agent workflows.
ccs/Qwen/Qwen3-Embedding-8B Embeddings only. Use /v1/embeddings, and set encoding_format: "float" or the request is rejected.
ccs/Qwen/Qwen3-VL-8B-Instruct-FP8 Vision-capable. Accepts image_url content parts as base64 data URLs.
ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink Identical to ccs/Qwen/Qwen3.6-35B-A3B-FP8 apart from chat_template_kwargs.enable_thinking = false, which suppresses the thinking phase. Preferred for latency-sensitive, non-tool calls. Not recommended for multi-tool agent loops: in testing it skipped a tool call and fabricated the answer in about one in five multi-tool turns (the thinking variant: none). Use ccs/Qwen/Qwen3.6-35B-A3B-FP8 for tool-calling.

This table covers quirks we have hit in practice; it is not a complete capability matrix. Anything not listed behaves the way the model and the inference gateway behave; this portal adds nothing.

Currently Available

ccs/deepseek-r1:32b ccs/gpt-oss:120b ccs/llama3.1:8b ccs/Qwen/Qwen3-Embedding-8B ccs/Qwen/Qwen3-VL-8B-Instruct-FP8 ccs/Qwen/Qwen3.6-35B-A3B-FP8 ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink ccs/qwen3:32b

Rendered live from the gateway each time this page loads. This is the catalog as of right now, not a fixed list.

To list models programmatically:

curl https://llm.ccs.uky.edu/v1/models -H "Authorization: Bearer sk-YOUR_KEY_HERE"

Troubleshooting

Errors raised here (auth, key permissions, upstream connectivity) come back as OpenAI-style error JSON. Errors raised by the model or the inference gateway pass through unchanged, so their wording is whatever upstream sent.

401 Unauthorized - "Missing API key" / "Invalid or expired API key"
No key was sent, or the key is wrong, expired, or revoked. Send it as Authorization: Bearer <key> or x-api-key: <key>. Personal and project keys are managed in Locksmith, which shows whether a key has expired or been revoked; a cluster job key stops working as soon as its SLURM job ends.
403 - "key not allowed to access model"
This key cannot use that model. Commercial cloud models through Azure (azure/...) need your project key; local models (ccs/...) work on your personal key. The message lists the model groups the key you sent can reach. A cluster job key can never reach the cloud models. Switch keys; asking CCS to widen a key is not the fix.
400 - rejected by the model
The inference gateway or the model refused the request body. The most common causes here are sending tools to ccs/deepseek-r1:32b, which does not support tool calling, and omitting encoding_format: "float" on an embeddings call. The response body carries the upstream message. Read it.
404 - unknown model
The model id does not exist in the current catalog. Ids are case-sensitive and include slashes and colons. Call GET /v1/models for the live list -- models come and go as GPU hosts are re-provisioned, so an id that worked last month may be gone.
429 - rate_limit_error, or the model is busy
Too fast, or too much at once (see Is there a rate limit?), or every worker is busy generating. Either way the response carries a Retry-After header, which the official OpenAI and Anthropic SDKs honour automatically; if you wrote your own client, back off for that many seconds and retry. Waiting fixes this one.
429 - budget_exceeded
This project's cloud budget is spent. Local models still work on this key. Contact your project PI to add funds.

The body says "Budget has been exceeded! ... Current cost: ..., Max budget: ..." and the response carries x-should-retry: false and no Retry-After: waiting does not fix this one, so stop retrying. The same message, without a Team= name, means your share of the project's pool is used up rather than the whole pool; the PI sets that cap too. Pools are one-off grants unless the PI arranged a periodic reset; the PI sees the balance and any reset date in Locksmith. Only cloud calls are refused - ccs/... models keep answering on the very same key.
502 - "Upstream gateway connection failed"
This host could not reach the inference gateway. Check https://llm.ccs.uky.edu/health/readiness. If checks.upstream is false, the outage is upstream rather than yours. Retry in a few minutes.
504 - "Upstream gateway timeout"
No response within the 10-minute read timeout. Reduce max_tokens, pick a smaller model, or stream. Streaming also stops intermediaries idling the connection out.
503 - "Gateway not configured"
The service is missing its upstream credential. That is an operator problem, not yours. Contact CCS.
Connection refused / cannot reach server
Check https://llm.ccs.uky.edu/health/liveness from your network - it should return 200. This address is reachable from the public internet, so no VPN is required and it works from off campus.

If that succeeds but your request still fails, the problem is the request rather than the network - check the API path and your key.

Known Limitations

This portal does not implement the API. It authenticates you and forwards the request verbatim to the CCS inference gateway, which serves the models. So what works is a property of the model you chose and the inference gateway, not of anything here. We add no samplers, no defaults, and no model filtering, and we do not rewrite responses.

The one exception: on streaming OpenAI-style requests we set stream_options.include_usage if you did not, so the final SSE chunk carries real token counts for usage accounting. If you already set it, yours is used.

The table below reflects what we have observed in practice. Treat it as guidance, not a contract. Verify against the model you actually intend to use.

FeatureStatusNotes
Chat completionsWorksStreaming and non-streaming
EmbeddingsWorksPass encoding_format: "float" explicitly; dimensions is ignored
Streaming (SSE)WorksProxied with buffering disabled, so chunks arrive as the model produces them
temperature, top_p, max_tokens, stopWorksStandard OpenAI sampling fields, honoured by the models in the catalog
Vision (images)Model-dependentWorks with a vision-capable model such as ccs/Qwen/Qwen3-VL-8B-Instruct-FP8. Base64 data URLs; we do not fetch remote image URLs on your behalf.
Tool / function callingModel-dependentStandard tools / tool_choice. Quality varies sharply by model; the -nothink variants and ccs/qwen3:32b are the practical choices. ccs/deepseek-r1:32b rejects tools with 400.
Thinking / reasoningModel-dependentReasoning models expose their chain of thought as a reasoning field, a delta.reasoning when streaming, or inline in the text, depending on the model. Not every client renders it. Use a -nothink variant to switch it off.
frequency_penalty, presence_penalty, seed, logprobs, response_formatVariesForwarded untouched. Whether a given model and serving engine honours, ignores, or rejects each one is decided upstream. Test before relying on it.
Engine-native sampler params (top_k, min_p, repeat_penalty, ...)VariesNot part of the OpenAI spec. Forwarded as sent; acceptance depends entirely on the upstream serving engine. There is no alternative native endpoint exposed through this host.
Anthropic Messages APIWorks/v1/messages is served upstream for Claude Code and the Anthropic SDKs: text, vision, streaming and tool use. Anthropic's signed thinking blocks cannot be reproduced by non-Claude models.
Model catalogChangesModels are added and removed as GPU hosts are re-provisioned. Query GET /v1/models; do not hard-code ids.

If a parameter does not behave as you expect, the place to look is the model and the inference gateway. Report it to CCS with the model id and the exact request body and we will chase it upstream.

Frequently Asked Questions

Are my prompts and responses logged?
Not by this host. We record metadata only - timestamp, model, status code, token counts, latency, client address, and the last four characters of the key; prompt and response content are never stored.
Where do I get an API key?
From locksmith.ccs.uky.edu. This portal does not issue personal keys. Keys for OOD cluster jobs are minted automatically by the job and are bound to it.
How many API keys can I have, and how long do they last?
Both are set by Locksmith policy. Manage, extend, and revoke your keys there.
What happens when my key expires?
Requests return 401 Unauthorized. Renew or reissue in Locksmith.
Can I use this from off-campus?
Yes. https://llm.ccs.uky.edu is reachable from the public internet and no VPN is needed. Your API key is what authenticates you, not your network location.

The one exception is cluster job provisioning at /api/v1/*, which stays campus-only on llm-internal.ccs.uky.edu. That is called by SLURM jobs on LCC / MCC / ECC, not by people.
What models are available?
Call GET /v1/models, or check the Status Page. The catalog changes as GPU hosts are re-provisioned, so query it rather than hard-coding ids.
Can I request a specific model?
Contact CCS. Models run on CCS GPU servers, so whether a given model can be added depends on GPU memory, the serving engine, and what else is deployed. It is a capacity conversation, not a download.
Is there a rate limit?
Yes, per API key, in two forms: a limit on how many requests you can start per second, and a limit on how many you can have in flight at once. There is no cap on total requests per day and no token quota.

Both are set well above what the GPUs can actually serve, so they exist to stop a runaway script, not to pace normal work. In practice you will not reach them.

The exact figures are deliberately not published here: they are tuned to the hardware behind the service and move as capacity is added. What stays constant is the contract -- if you hit a limit you get a 429 with a Retry-After header, and honouring it is all you need to do. If you have a workload that genuinely needs more than one key allows, contact CCS rather than working around it with extra keys: the service is shared, and the limits are what keep it answering for everyone.
I got a 429. Am I being rate limited?
Read the type in the body first. If it says budget_exceeded, this is not about speed at all: a project's cloud pool is spent - see the budget 429, and do not retry. Otherwise a 429 means one of two quite different things:
  • You are sending too fast - starting requests faster, or holding more in flight at once, than one key is allowed. Under your control: slow down, or run fewer in parallel.
  • The model is busy - every request is holding a worker while it generates. Not under your control: the fix is to wait, reduce how many requests you run in parallel, or use a model that is already loaded. See the status page for what is running now.
The second is far more common. If you are running a handful of requests at a time and seeing 429, it is capacity, and sending them more slowly one at a time will help where sending them all at once will not.

Either way the response carries a Retry-After header. The official OpenAI and Anthropic SDKs honour it automatically; if you wrote your own client, wait that many seconds before retrying. Please be mindful of shared resources.
Can I call the API directly from a web page?
No, and this is deliberate rather than an oversight. A browser will refuse a cross-origin request to /v1 from another site.

The reason is that to make such a call, the page has to carry your API key in its JavaScript, where anyone who opens the developer tools can read it. Your key identifies you personally, so publishing it is not a small mistake.

Call the API from your server instead and have your page talk to that. Your key stays on a machine you control, and it is what the official OpenAI and Anthropic SDKs assume you are doing.

The one exception is our own test page, which works because it is served from this site - a page calling back to its own origin never involves cross-origin rules at all.
Who can I contact for help?
Email the CCS team or visit ccs.uky.edu for support contacts.