Documentation
How to call the CCS AI Inference Gateway from a shell, a notebook, an editor or a batch job.
Set up your tool
Pick what you are using. Every one of them needs the same two
things: the base URL https://llm.ccs.uky.edu/v1 and your key.
- OpenWebUI
- Jupyter AI
- VS Code / Continue
- Cursor
- Claude Code
- Python
- TypeScript
- curl
- Jupyter notebook
- LangChain
- LlamaIndex
- Embeddings / RAG
An empty reply is usually not an outage
Reasoning models spend part of your max_tokens budget thinking before they answer. If the budget runs out while the model is still thinking, you get back an empty reply, HTTP 200 with no error. The tell is completion_tokens coming back exactly equal to the max_tokens you asked for. There is no single safe number to set: try about 2,000 for a simple question and 8,000 or more for multi-step work, then raise it and retry if you see that signature. ccs/gpt-oss:120b is barely affected. On ccs/Qwen/Qwen3.6-35B-A3B-FP8 (thinking) the reply is never empty; instead the thinking text is cut off with no answer. Non-thinking and plain models are not affected. For multi-tool or agent work, prefer the thinking alias. The status page marks which models think first.
Which model should I use?
Several models are available, and any of them will answer an everyday request well: a question, a conversation, a page or two of text, a snippet of code. The choice below only really matters when your document is long, or when you are doing something specific like reading an image or building a search tool.
Everything in this section was checked with real requests on 2026-08-17. None of it is copied from a manufacturer's sheet.
The 30-second answer
| If you want to… | Use |
|---|---|
| Work with a long document or a large amount of code | Qwen3.6 35B (no-think)ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink |
| Build a tool-using assistant or agent (function/tool calls) | Qwen3.6 35B (thinking)ccs/Qwen/Qwen3.6-35B-A3B-FP8 |
| Automate a plain, non-tool task (summarise, extract, transform) | Qwen3.6 35B (no-think) |
| Ask a quick question and get a fast answer | Llama 3.1 8Bccs/llama3.1:8b |
| See careful step-by-step reasoning you can read | Qwen3.6 35B (thinking)ccs/Qwen/Qwen3.6-35B-A3B-FP8 |
| Ask about a picture, screenshot or chart | Qwen3-VL 8Bccs/Qwen/Qwen3-VL-8B-Instruct-FP8 |
| Build a search tool over your own documents | Qwen3-Embedding 8Bccs/Qwen/Qwen3-Embedding-8B |
Qwen3.6 35B (no-think): "no-think" gives you
just the answer, with no visible deliberation. It is not a weaker
model; it is the same model with the step-by-step reasoning turned
off so it replies faster, and it is the recommended default for chat, long documents and code. For a tool-calling agent, prefer the thinking variant (ccs/Qwen/Qwen3.6-35B-A3B-FP8): without the reasoning phase, the no-think alias can silently skip a tool call and answer from guesswork.
How much can I paste?
Every model has a limit on how much it can take in one request, and that limit counts your input only: every message in the conversation added together, including the system prompt and anything your program retrieved and pasted in. The length of the reply you ask for does not count towards it. We give the limit in pages of ordinary double-spaced text, because that is what you actually paste. These are rough equivalents, not exact conversions. The real figure depends on formatting and whether it is prose or code.
| Limit | Roughly the size of… | If you are pasting code |
|---|---|---|
| about 25 pages (approximately 15,500 tokens) | one journal article, a long email thread, a 10,000-word report | about 1,000–1,500 lines |
| about 50 pages (approximately 33,000 tokens) | two or three journal articles, a grant proposal | about 2,000–3,000 lines |
| about 100 pages (approximately 65,000 tokens) | a Master's thesis, a short book, a full literature review | about 4,000–6,000 lines |
| about 200 pages (approximately 131,000 tokens) | a PhD dissertation, a full manual | about 8,000–12,000 lines |
The one thing everyone needs to know
Every model tells you when your text is too long. In almost every case you get a clear error naming the limit, nothing is sent, and you never get an answer built from a partial document. One narrow exception is noted below.
Two things follow from that. First, the limit counts your input only, not the reply you ask for. Lowering max_tokens will not bring a long document under the limit; only shortening the input will. Second, you no longer have to remember
which models are “safe”: pick one whose limit fits
your text, and if you go over you simply get an error instead
of a wrong answer.
Rule of thumb: pick a model whose limit fits what you are pasting. If it is longer than about 25 pages, use Qwen3.6 35B (no-think). It holds about 200 pages (approximately 131,000 tokens), the most of anything here.
It adds up across a conversation. The limit applies to everything you send in one request, and a chat client resends the whole conversation every turn. System prompt, any text your program retrieved and pasted in, every earlier question and answer, and the new question are all counted together. This is why a chat that worked on turn 1 can start failing on turn 6 with nothing changed: the conversation itself grew past the limit. If that happens, start a new conversation or move to a model with a larger limit.
One narrow exception. For unusually dense input such as hashes, base64, or minified data, a prompt just under the limit can occasionally be trimmed with no error. If usage.prompt_tokens comes back at exactly the model’s limit, that happened and the front of your input was dropped. This matters most when the text you retrieve is itself dense (source code, logs, CSV, JSON): if you are building search over material like that, prefer a model with a much larger limit so you are never near the edge.
How much each model holds
These limits were re-measured with real requests on 2026-08-17. Every model now refuses over-length input with a clear error, so the only question is whether a model's limit is big enough for what you are pasting.
| Model | You can paste | If you go over |
|---|---|---|
| Qwen3.6 35B (both versions) | about 200 pages (approximately 131,000 tokens) | Clear error. Nothing lost. |
| gpt-oss 120B | about 190 pages (approximately 124,000 tokens) | Clear error. Nothing lost. |
| Qwen3-VL 8B (reads images too) | about 100 pages (approximately 65,000 tokens) | Clear error. Nothing lost. |
| Qwen3-Embedding 8B | about 50 pages per item (approximately 33,000 tokens) | Clear error. Nothing lost. |
| DeepSeek-R1 32B | about 45 pages (approximately 28,000 tokens) | Clear error. Nothing lost. |
| Qwen3 32B | about 25 pages (approximately 15,500 tokens) | Clear error. Nothing lost. |
| Llama 3.1 8B | about 25 pages (approximately 15,500 tokens) | Clear error. Nothing lost. |
DeepSeek-R1 32B is a special case. Its limit of about 45 pages is not a hardware ceiling but a reliability one: past roughly that size it can no longer reliably find a fact in what you gave it, and may answer “not present” with full confidence. For a long document, use Qwen3.6 35B (either version) or gpt-oss 120B instead.
A note on each model
| Model | Good for | You can paste | If you go over |
|---|---|---|---|
Qwen3.6 35B (no-think)ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink | The default for chat, long documents and code. Very fast, and gives you the answer and nothing else. For tool-calling agents, use the thinking variant instead. | about 200 pages (approximately 131,000 tokens) | Clear error. Nothing lost. |
Qwen3.6 35B (thinking)ccs/Qwen/Qwen3.6-35B-A3B-FP8 | The same model with its reasoning shown. Use it when you want to read how it reached the answer. | about 200 pages (approximately 131,000 tokens) | Clear error. Nothing lost. |
Qwen3-VL 8Bccs/Qwen/Qwen3-VL-8B-Instruct-FP8 | The only model that can look at a picture: screenshots, charts, scanned pages, diagrams, photos. | about 100 pages (approximately 65,000 tokens) | Clear error. Nothing lost. |
gpt-oss 120Bccs/gpt-oss:120b | The largest model here. General chat and demanding work when quality matters more than speed. | about 190 pages (approximately 124,000 tokens) | Clear error. Nothing lost. |
DeepSeek-R1 32Bccs/deepseek-r1:32b | Problems that need working through step by step. Reliably handles about 45 pages; for a longer document use Qwen3.6 35B or gpt-oss 120B. Cannot do tool calling: sending tools returns an error, so it cannot be used by an agent. | about 45 pages (approximately 28,000 tokens) | Clear error. Nothing lost. |
Qwen3 32Bccs/qwen3:32b | Reasoning on shorter inputs. | about 25 pages (approximately 15,500 tokens) | Clear error. Nothing lost. |
Llama 3.1 8Bccs/llama3.1:8b | Quick questions, short summaries, rewriting and drafting. Usually replies in about a second. | about 25 pages (approximately 15,500 tokens) | Clear error. Nothing lost. |
Qwen3-Embedding 8Bccs/Qwen/Qwen3-Embedding-8B | Building search over your own documents. It turns text into numbers so you can search by meaning; it does not chat. | about 50 pages per item (approximately 33,000 tokens) | Clear error. Nothing lost. |
A few quirks worth knowing
- Llama 3.1 8B and search. It will happily accept a “turn this text into search numbers” request and hand back numbers that look real but are not. A search tool built on them quietly gives poor results, and nothing in the reply reveals the problem. Send search and “embedding” requests only to Qwen3-Embedding 8B.
- Qwen3.6 35B (no-think) with automated tools. When you offer it several tools at once, it sometimes answers from memory instead of using one, and can state a made-up figure with full confidence. If your program relies on a tool to fetch live information, do not treat “no tool used” as “nothing was needed”; asking again and requiring the tool fixed it every time we tried.
- Very short replies come back empty. The thinking version of Qwen3.6 35B and Qwen3 32B can return nothing at all if you ask for a very short reply, because they spend the whole budget thinking before they start. There is no single safe number: try about 2,000 tokens for a simple reply and 8,000 or more for multi-step work, then retry if it comes back empty. Or use a model that does not think first (though for multi-tool or agent work, prefer the thinking alias).
- Qwen3.6 35B (thinking) is slower, because
it reasons before it answers. As of 2026-08-28 it
keeps that reasoning in a separate
reasoning_contentfield andcontentholds only the answer; until that date the reasoning bled into the answer text and you could see a stray</think>in what came back. If you wrote code to strip that, it is no longer needed. The no-think version returns just the answer with no reasoning at all, which is another reason it is the default for scripts and automation. - Qwen3-VL 8B and Qwen3-Embedding 8B were the cleanest in testing: no problems found.
Writing code against the API? The service measures length in small units rather than pages: roughly 650 to a page, so the “about 25 pages” limit is about 15,500 of them. You never need to work this out; paste length is what matters.
Last tested 2026-08-17. Every limit and behaviour above was verified with real requests, not taken from documentation.
Overview
The CCS AI Inference Gateway provides access to large language models, hosted on CCS GPU hardware and, on a project key, commercial cloud models through Azure, all through one OpenAI-compatible API. This is the same API format used by OpenAI, and it is the de facto industry standard for LLM APIs.
Note: Access is granted per person on request. If you do not have an API key yet, ask for one at Locksmith. The service is in testing, so expect occasional interruptions.
How the pieces fit together
This host is a thin authenticating proxy. It hosts no models and
performs no format translation. Your request is authenticated against your API key,
then forwarded verbatim to the CCS inference gateway, which runs the
local models on CCS GPU servers and relays project-key requests for
azure/ models to the cloud. The response is streamed back to you unchanged.
Two consequences worth knowing:
- API behaviour is defined upstream, not here. Which sampling parameters, tool-calling features, and response fields you get for a given model is decided by the inference gateway and the model itself. We do not add, drop, or rewrite parameters.
- The model catalog changes. Models are added and removed as GPU
hosts are re-provisioned. Always discover the current list with
GET /v1/modelsrather than hard-coding names.
What does "OpenAI-compatible" mean for you?
It means any tool, library, or application that works with the OpenAI API works with our gateway. You just change two things:
- The base URL to
https://llm.ccs.uky.edu/v1 - The API key to a key from Locksmith - your personal key, or a project key if you want that project's cloud models (see API keys)
No code changes, no special libraries, no adapter layers. The following tools and frameworks work out of the box:
Base URL: https://llm.ccs.uky.edu/v1
Privacy: Your prompts and responses are never logged. We record metadata only - timestamp, model, status code, token counts, latency, client address, and the last four characters of the key; prompt and response content are never stored.
Getting Started
- Get a key from Locksmith - API keys are issued by
CCS Key Management at locksmith.ccs.uky.edu.
Sign in there with your institutional credentials via CILogon (University
of Kentucky, ACCESS-CI, or another participating institution), open
AI Inference Keys, and create your
personal key. If you belong to a research project with a commercial-cloud
budget, the same page also shows a project key for each such project.
All keys issued by Locksmith begin with
sk-. - Copy your key - The full key is shown only once. Copy it immediately and store it securely.
- Point your client here - Set the base URL to
https://llm.ccs.uky.edu/v1and use your key as the API key. See the integration guides below. - Discover the models - Call
GET /v1/models(or check the Status Page) for the current catalog, then pass one of those ids asmodel.
Running an LLM from a cluster batch job (OOD "Launch an LLM") is
different: those jobs are issued a short-lived ccs-llm- key automatically
for the life of the SLURM job. You do not request those by hand. See
API Key Management.
API Key Management
There are three kinds of key. All are sent the same way and work on every endpoint documented here; they differ in which models they can call and who pays. The key you configure in a tool decides both - there is no header, model-name prefix or setting to choose a project; the key is the choice.
| Personal key | Project key | Cluster job key | |
|---|---|---|---|
| Issued by | Locksmith, one per person | Locksmith, one per project you belong to | Automatically, by the OOD "Launch an LLM" job when it starts |
| Format | sk-<random> |
sk-<random> |
ccs-llm-<random> |
| Models | Local models (ccs/...) |
Local models and commercial cloud models through Azure (azure/...) |
Local models only |
| Who pays | Nobody - local models are free | Cloud calls draw on that project's dollar pool; local calls are free | Nobody |
| Who it is for | You: laptops, notebooks, editors, chat UIs | You, when working for that project | A single SLURM job on LCC / MCC / ECC |
| Lifetime | Set by Locksmith policy when issued | While you are a member of the project | Bound to the job: revoked when the job ends, and swept by a reaper as a backstop |
| Revoke | In Locksmith. Takes effect within about a minute; cannot be undone. | In Locksmith, by you or the project PI. Takes effect within about a minute. | Automatic at job exit; a CCS admin can also kill it immediately |
| Shown once | The full key is displayed only at creation. Only a hash is stored, never the key itself. | ||
This portal does not mint personal or project keys. Both come from
Locksmith. Examples throughout this page use sk-YOUR_KEY_HERE as the
placeholder for whichever sk- key you chose; if you are inside a cluster job,
substitute the ccs-llm- key the job was given and everything else is identical.
Commercial cloud models through Azure cost money. Each project has a dollar pool that its PI arranged; every cloud call on a project key draws on it, and when the pool is spent, cloud calls on that key are refused until the PI adds funds (pools do not refill by themselves unless the PI set that up). Local models keep working on the same key throughout. The PI sees the balance in Locksmith. Cluster job keys can never reach the cloud models, by design: put your project key in the job if a batch job needs one.
Security: API keys are personal. Do not share them, commit them to git, or embed them in client-side code. Use environment variables or secret managers.
OpenAI API Compatibility
Our API implements the OpenAI chat completions specification. Any code written for the OpenAI API works by changing just the base URL and API key:
# Before: talking to OpenAI
client = OpenAI(api_key="<your OpenAI key>")
# After: talking to the CCS AI Inference Gateway - add base_url, swap the key
client = OpenAI(
base_url="https://llm.ccs.uky.edu/v1",
api_key="sk-YOUR_KEY" # issued by locksmith.ccs.uky.edu
)
Everything else stays the same: client.chat.completions.create(),
client.models.list(), streaming, etc. The only other change you will
make is the model argument, which must name a model from our catalog
(see Available Models) rather than an OpenAI model.
Note: CCS keys and OpenAI keys share the sk- prefix
but are unrelated. A CCS key only works against
https://llm.ccs.uky.edu.
API Reference
Base URL: https://llm.ccs.uky.edu/v1
Authentication: Authorization: Bearer sk-YOUR_KEY
or x-api-key: sk-YOUR_KEY. Both headers are
accepted on every endpoint, and any one key (personal or project) works for both the
OpenAI-compatible (/v1/chat/completions) and
Anthropic-compatible (/v1/messages) routes.
| Method | Endpoint | Description | Auth |
|---|---|---|---|
POST | /v1/chat/completions | Chat completions (primary endpoint) | Required |
POST | /v1/messages | Anthropic-compatible Messages API (see Claude Code / Anthropic SDK) | Required |
POST | /v1/completions | Legacy text completions | Required |
POST | /v1/embeddings | Generate embeddings | Required |
GET | /v1/models | List available models | Required |
GET | /health/liveness | Is the service process up | None |
GET | /health/readiness | Is the service up and the upstream gateway reachable (not per-model health; see the status page for that) | None |
Chat Completions Request Body
| Parameter | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Model ID (e.g., ccs/llama3.1:8b) |
messages | array | Yes | Array of {role, content} objects |
temperature | float | No | Sampling temperature (0-2) |
max_tokens | integer | No | Maximum tokens to generate |
top_p | float | No | Nucleus sampling threshold |
stream | boolean | No | Enable Server-Sent Events streaming |
stop | array | No | Stop sequences |
tools, tool_choice | array / string | No | Function calling. Not every model accepts these; see Available Models |
This table lists the parameters in common use. The gateway does not filter the request body: any other OpenAI field you send is forwarded as-is to the inference gateway, which decides whether the chosen model honours it, ignores it, or rejects the request. See Known Limitations.
Response Format
{
"id": "chatcmpl-1712345678000",
"object": "chat.completion",
"created": 1712345678,
"model": "ccs/llama3.1:8b",
"choices": [{
"index": 0,
"message": {"role": "assistant", "content": "..."},
"finish_reason": "stop"
}],
"usage": {
"prompt_tokens": 25,
"completion_tokens": 150,
"total_tokens": 175
}
}
Embeddings
Embedding models convert text into numerical vectors for use in search, clustering, RAG (retrieval-augmented generation), and similarity tasks. The endpoint is POST /v1/embeddings, same format as OpenAI.
Pass encoding_format explicitly
Send "encoding_format": "float" on every request to
ccs/Qwen/Qwen3-Embedding-8B. An earlier version of the
service rejected requests that left it out; that has since been
relaxed, but the OpenAI SDKs default to base64, so passing it
explicitly is what makes your code do the same thing on every
client and across service upgrades. It is always safe to send, and
every example below does.
dimensions is ignored. Ask for 1,024
and you still get 4,096 back, with no error and no warning. If you
need shorter vectors, truncate them yourself: these are Matryoshka
embeddings, so the first N numbers of a vector are a valid shorter
vector. Truncate every vector in an index to the same length, and
re-normalise after truncating.
This model is embeddings only. Sending it to
/v1/chat/completions will fail.
Request
curl -X POST https://llm.ccs.uky.edu/v1/embeddings \
-H "Authorization: Bearer sk-YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "ccs/Qwen/Qwen3-Embedding-8B",
"input": "What is machine learning?",
"encoding_format": "float"
}'
The input field accepts either a single string or an array of strings. When you pass an array, the API returns one embedding per input in the same order, which is more efficient than making separate requests:
Keep a batch to 256 items or fewer. Very large float batches have been measured to take down a service worker, which returns errors to unrelated users, not just to you. When you are indexing a whole document collection, loop over batches of 256 rather than sending one enormous array.
curl -X POST https://llm.ccs.uky.edu/v1/embeddings \
-H "Authorization: Bearer sk-YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "ccs/Qwen/Qwen3-Embedding-8B",
"input": ["First document", "Second document", "Third document"],
"encoding_format": "float"
}'
Response
{
"object": "list",
"data": [{
"object": "embedding",
"embedding": [0.0023, -0.0091, 0.0152, ...],
"index": 0
}],
"model": "ccs/Qwen/Qwen3-Embedding-8B",
"usage": {"prompt_tokens": 6, "total_tokens": 6}
}
Python Example
from openai import OpenAI
client = OpenAI(
base_url="https://llm.ccs.uky.edu/v1",
api_key="sk-YOUR_KEY"
)
# Generate embedding for a single string
response = client.embeddings.create(
model="ccs/Qwen/Qwen3-Embedding-8B",
input="What is machine learning?",
encoding_format="float", # always pass this explicitly
)
vector = response.data[0].embedding
print(f"Dimensions: {len(vector)}")
print(f"First 5: {vector[:5]}")
# Generate embeddings for multiple strings at once
response = client.embeddings.create(
model="ccs/Qwen/Qwen3-Embedding-8B",
input=["First document", "Second document", "Third document"],
encoding_format="float",
)
for item in response.data:
print(f"Index {item.index}: {len(item.embedding)} dimensions")
# Compare similarity of two texts
import numpy as np
def cosine_sim(a, b):
return np.dot(a, b) / (np.linalg.norm(a) * np.linalg.norm(b))
def embed(text):
r = client.embeddings.create(
model="ccs/Qwen/Qwen3-Embedding-8B",
input=text,
encoding_format="float",
)
return r.data[0].embedding
cat, dog, plane = embed("cat"), embed("dog"), embed("airplane")
print(f"cat vs dog: {cosine_sim(cat, dog):.4f}")
print(f"cat vs airplane: {cosine_sim(cat, plane):.4f}")
# Semantically related pairs score higher than unrelated ones.
Node.js Example
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://llm.ccs.uky.edu/v1",
apiKey: "sk-YOUR_KEY",
});
// Single string
const single = await client.embeddings.create({
model: "ccs/Qwen/Qwen3-Embedding-8B",
input: "What is machine learning?",
encoding_format: "float", // always pass this explicitly
});
console.log(`Dimensions: ${single.data[0].embedding.length}`);
// Array of strings (batch)
const batch = await client.embeddings.create({
model: "ccs/Qwen/Qwen3-Embedding-8B",
input: ["First document", "Second document", "Third document"],
encoding_format: "float",
});
for (const item of batch.data) {
console.log(`Index ${item.index}: ${item.embedding.length} dimensions`);
}
Available Embedding Models
At the time of writing the catalog contains a single embedding model,
ccs/Qwen/Qwen3-Embedding-8B. That can change: call
GET /v1/models or check the
Status Page, where embedding models are
labelled "Embedding", for the current list.
Do not hard-code the vector length. Read it from the response
(len(response.data[0].embedding)) so your index does not silently
break if the hosted model changes.
Note: Embedding models use /v1/embeddings, not /v1/chat/completions. They return vectors, not text. Point an OpenAI-compatible embeddings client at /v1/embeddings to test embedding models.
Set the embedding model deliberately, then check it
A chat model asked for an embedding does not refuse. It returns
a vector of a plausible length, correctly normalised, echoing back
whatever model name you asked for. ccs/llama3.1:8b
returns 4096 numbers, exactly like the real embedder does, so
nothing in the response tells you the wrong model answered.
The failure that follows is quiet: a search index built on the wrong vectors still returns results, still looks fine on a handful of easy test documents, and simply retrieves the wrong document some of the time in production. A smoke test will not catch it.
So set the embedding model once, explicitly, in one place in your
code, and assert on it. Send embedding requests
only to ccs/Qwen/Qwen3-Embedding-8B.
If you ever re-index, re-embed everything with the same model:
vectors from two different models cannot be compared.
Budgeting context for RAG and web search
If you are building retrieval-augmented generation, or letting a model read web pages, the model's limit is the constraint that will bite you first, and it will bite in a confusing way. What you send each turn is not just the question:
system prompt + (chunks retrieved x tokens per chunk) + conversation so far + the question
All of it counts, and all of it is input, so lowering
max_tokens does nothing. The conversation grows every
turn, which is why a RAG chat frequently works perfectly on the first
question and starts returning errors several questions later.
Working the arithmetic against the published limits, taking about 600 tokens of system prompt and question as overhead:
| Retrieval settings | Roughly what you send | Where it fits |
|---|---|---|
| 4 chunks of 512 tokens | about 2,650 | Every model here. |
| 8 chunks of 1,024 tokens | about 8,800 | Every model here. |
| 8 chunks of 1,024, plus about five turns of conversation | about 12,800 | Fits the 25-page models on the first question, then fails as the conversation grows. Comfortable from DeepSeek-R1 32B upwards. |
| 16 chunks of 1,024 tokens | about 17,000 | Over the limit on Qwen3 32B and Llama 3.1 8B. Fits DeepSeek-R1 32B and larger. |
| 20 chunks of 1,500 tokens | about 30,600 | Needs Qwen3-VL 8B or larger; comfortable only on gpt-oss 120B and Qwen3.6 35B. |
The simple answer: point retrieval and web-search traffic at
ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink (about 200 pages,
approximately 131,000 tokens) or ccs/gpt-oss:120b (about
190 pages, approximately 124,000 tokens). None of the settings
above can overrun either one, and choosing one of them removes this
whole class of failure rather than tuning around it. Reach for a
smaller model for retrieval only if you have measured that you need
its speed.
If you must stay on a smaller model, the levers are: retrieve fewer chunks, make the chunks shorter, or drop older turns from the conversation you resend. Those are the only three that work.
Do not treat these limits as permanent. They change as models are re-provisioned, so check the status page, which lists the current limit for every model and needs no key, before you settle on a chunk count for a long-lived index.
Settings for chat apps (Open WebUI and similar)
If you run your own chat application against this service, these are the settings to put in it. The defaults that ship with these apps are written for hosted commercial services with much larger limits, and leaving them alone is the usual cause of errors that look like the service is broken when it is not. A few of them can also degrade the service for other people, so they are worth setting deliberately.
| Setting | Use this | Why |
|---|---|---|
| API base URL | https://llm.ccs.uky.edu/v1 | The only address to use. Everything goes through here. |
| API key | Your own key | One key identifies one person. Do not share it or put it in a page others can read. |
| Chat model | ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink | Fast, and its limit is the largest here, so long chats do not hit it. |
| Embedding model | ccs/Qwen/Qwen3-Embedding-8B | The only real one. See the warning below; this is the setting people get wrong. |
| Embedding batch size | 256 or fewer | Larger batches can take down a service worker and return errors to other users. |
| Chunk size | 500–1,000 tokens | Ordinary RAG chunk sizes. The hard ceiling is 32,768 tokens per item. |
| Chunks retrieved (top-k) | 4–8 to start | Every retrieved chunk counts against the model's input limit. See the budget table. |
| Embedding dimensions | Leave unset | The setting is ignored. You always get 4,096 back whatever you ask for. |
The one setting that fails silently: the embedding model
If you type a chat model into the embedding field, you get no error. Measured on this service:
Model sent to /v1/embeddings | What comes back |
|---|---|
ccs/Qwen/Qwen3-Embedding-8B | 4,096 numbers. Correct. |
ccs/llama3.1:8b | 4,096 numbers, same shape, echoes the name you sent. Wrong, and undetectable. |
ccs/deepseek-r1:32b | 5,120 numbers. Wrong. |
The llama3.1:8b case is the dangerous one: it is the same
length and the same shape as the real thing, so nothing in the reply
tells you it is wrong. Your document search will build, run, and
return plausible results while retrieving the wrong document a large
fraction of the time. A quick test on a handful of documents will
pass.
Set it once, deliberately, and check it. If you change it later, re-index everything: vectors made by two different models cannot be compared.
If your chat starts failing after a few messages
That is the input limit, not a fault. Your app resends the whole conversation every turn, so a chat that worked at the start can grow past the limit later, especially with documents attached. Start a new conversation, or use a model with a larger limit. Lowering the maximum reply length does not help, because the limit counts only what you send. See how much each model holds.
Python / OpenAI SDK
# pip install openai
from openai import OpenAI
client = OpenAI(
base_url="https://llm.ccs.uky.edu/v1",
api_key="sk-YOUR_KEY_HERE"
)
# Simple chat
response = client.chat.completions.create(
model="ccs/llama3.1:8b",
messages=[
{"role": "system", "content": "You are a helpful research assistant."},
{"role": "user", "content": "What is gradient descent?"}
],
temperature=0.7,
max_tokens=500
)
print(response.choices[0].message.content)
# Output: "Gradient descent is an optimization algorithm used to minimize
# a function by iteratively moving in the direction of steepest
# descent as defined by the negative of the gradient..."
print(f"Tokens used: {response.usage.total_tokens}")
# Output: Tokens used: 142
# List all available models
for model in client.models.list().data:
print(model.id)
# Output:
# ccs/deepseek-r1:32b
# ccs/llama3.1:8b
# ccs/qwen3:32b
# ccs/Qwen/Qwen3-Embedding-8B
# ...
Using Environment Variables (recommended)
# Set in your shell profile or .env file:
# export OPENAI_API_KEY="sk-YOUR_KEY_HERE"
# export OPENAI_BASE_URL="https://llm.ccs.uky.edu/v1"
# Then in Python - no credentials in code:
client = OpenAI() # reads from environment automatically
response = client.chat.completions.create(
model="ccs/llama3.1:8b",
messages=[{"role": "user", "content": "Hello!"}]
)
TypeScript / Node.js
The official OpenAI Node.js / TypeScript SDK works with the CCS AI Inference Gateway. Install it with npm or your preferred package manager:
npm install openai
Chat Completion
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://llm.ccs.uky.edu/v1",
apiKey: "sk-YOUR_KEY_HERE",
});
async function main() {
const response = await client.chat.completions.create({
model: "ccs/llama3.1:8b",
messages: [
{ role: "system", content: "You are a helpful research assistant." },
{ role: "user", content: "What is gradient descent?" },
],
temperature: 0.7,
max_tokens: 500,
});
console.log(response.choices[0].message.content);
console.log(`Tokens used: ${response.usage?.total_tokens}`);
}
main();
Streaming
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://llm.ccs.uky.edu/v1",
apiKey: "sk-YOUR_KEY_HERE",
});
async function streamChat() {
const stream = await client.chat.completions.create({
model: "ccs/llama3.1:8b",
messages: [{ role: "user", content: "Write a short essay on AI ethics." }],
stream: true,
});
for await (const chunk of stream) {
const content = chunk.choices[0]?.delta?.content;
if (content) {
process.stdout.write(content);
}
}
console.log(); // newline at end
}
streamChat();
Embeddings
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://llm.ccs.uky.edu/v1",
apiKey: "sk-YOUR_KEY_HERE",
});
async function embed() {
const response = await client.embeddings.create({
model: "ccs/Qwen/Qwen3-Embedding-8B",
input: "What is machine learning?",
encoding_format: "float",
});
const vector = response.data[0].embedding;
console.log(`Dimensions: ${vector.length}`);
console.log(`First 5: ${vector.slice(0, 5)}`);
}
embed();
TypeScript Type Hints
The SDK exports full type definitions. Use them for better autocompletion and type safety:
import OpenAI from "openai";
import type {
ChatCompletion,
ChatCompletionMessageParam,
ChatCompletionCreateParamsNonStreaming,
} from "openai/resources/chat/completions";
const client = new OpenAI({
baseURL: "https://llm.ccs.uky.edu/v1",
apiKey: "sk-YOUR_KEY_HERE",
});
// Typed message array
const messages: ChatCompletionMessageParam[] = [
{ role: "system", content: "You are a helpful assistant." },
{ role: "user", content: "Explain TypeScript generics." },
];
// Typed request parameters
const params: ChatCompletionCreateParamsNonStreaming = {
model: "ccs/llama3.1:8b",
messages,
temperature: 0.7,
max_tokens: 500,
};
async function typedChat(): Promise<string | null> {
const response: ChatCompletion = await client.chat.completions.create(params);
return response.choices[0].message.content;
}
typedChat().then(console.log);
Using Environment Variables (recommended)
# Set in your shell profile or .env file:
# export OPENAI_API_KEY="sk-YOUR_KEY_HERE"
# export OPENAI_BASE_URL="https://llm.ccs.uky.edu/v1"
# Then in your code - no credentials needed:
const client = new OpenAI(); // reads from environment automatically
curl Examples
Chat Completion
curl -X POST https://llm.ccs.uky.edu/v1/chat/completions \
-H "Authorization: Bearer sk-YOUR_KEY_HERE" \
-H "Content-Type: application/json" \
-d '{
"model": "ccs/llama3.1:8b",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is the capital of Kentucky?"}
]
}'
Response:
{
"id": "chatcmpl-123",
"object": "chat.completion",
"created": 1712345678,
"model": "ccs/llama3.1:8b",
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"content": "The capital of Kentucky is Frankfort."
},
"finish_reason": "stop"
}],
"usage": {
"prompt_tokens": 28,
"completion_tokens": 9,
"total_tokens": 37
}
}
Streaming
curl -N -X POST https://llm.ccs.uky.edu/v1/chat/completions \
-H "Authorization: Bearer sk-YOUR_KEY_HERE" \
-H "Content-Type: application/json" \
-d '{
"model": "ccs/llama3.1:8b",
"messages": [{"role": "user", "content": "Write a haiku about computing."}],
"stream": true
}'
Response (Server-Sent Events):
data: {"id":"chatcmpl-456","object":"chat.completion.chunk","created":1712345678,"model":"ccs/llama3.1:8b","choices":[{"index":0,"delta":{"role":"assistant","content":"Silicon"},"finish_reason":null}]}
data: {"id":"chatcmpl-456","object":"chat.completion.chunk","created":1712345678,"model":"ccs/llama3.1:8b","choices":[{"index":0,"delta":{"content":" dreams"},"finish_reason":null}]}
data: {"id":"chatcmpl-456","object":"chat.completion.chunk","created":1712345678,"model":"ccs/llama3.1:8b","choices":[{"index":0,"delta":{"content":" awake"},"finish_reason":null}]}
...
data: {"id":"chatcmpl-456","object":"chat.completion.chunk","created":1712345678,"model":"ccs/llama3.1:8b","choices":[{"index":0,"delta":{},"finish_reason":"stop"}],"usage":{"prompt_tokens":15,"completion_tokens":12,"total_tokens":27}}
data: [DONE]
List Models
curl https://llm.ccs.uky.edu/v1/models \
-H "Authorization: Bearer sk-YOUR_KEY_HERE"
Response:
{
"object": "list",
"data": [
{"id": "ccs/deepseek-r1:32b", "object": "model", "created": 1712345678, "owned_by": "ccs"},
{"id": "ccs/gpt-oss:120b", "object": "model", "created": 1712345678, "owned_by": "ccs"},
{"id": "ccs/llama3.1:8b", "object": "model", "created": 1712345678, "owned_by": "ccs"},
{"id": "ccs/qwen3:32b", "object": "model", "created": 1712345678, "owned_by": "ccs"}
]
}
Abridged. This endpoint is the authoritative list of what you
can pass as model, and it changes as GPU hosts are re-provisioned --
query it rather than copying ids out of this page.
Health Check (no key needed)
curl https://llm.ccs.uky.edu/health/readiness
Response:
{"checks": {"config": true, "db": true, "upstream": true}, "status": "ready", "upstream": {"db": "connected", "cache": null}}
Returns 200 with "status": "ready" when the
inference gateway is reachable, 503 with "status": "not_ready"
when it is not.
OpenWebUI
OpenWebUI is an open-source ChatGPT-like interface that supports custom OpenAI-compatible backends.
- Open your OpenWebUI instance
- Go to Admin Panel → Settings → Connections
- Under OpenAI API, click Add Connection
- Set URL to:
https://llm.ccs.uky.edu/v1 - Set API Key to: your key from Locksmith (personal, or a project key to see that project's cloud models)
- Click the refresh icon to verify the connection
- Available models will appear in the model dropdown when starting a new chat
Tip: If you run OpenWebUI via Docker, you can also set OPENAI_API_BASE_URL and OPENAI_API_KEY as environment variables.
Jupyter AI (Jupyternaut Chat)
Jupyter AI adds a chat interface (Jupyternaut) and inline completions to JupyterLab. Each user configures their own settings, no shared environment variables needed.
Step-by-Step Setup
- Open the Jupyternaut chat panel in JupyterLab (left sidebar)
- Click the gear icon to open settings
- Fill in the fields exactly as shown below
- Scroll to the bottom and click Save Changes
Language Model
| Field | Exact Value |
|---|---|
| Completion model | OpenAI (general interface) :: * (select from dropdown) |
| Model ID | ccs/llama3.1:8b (no trailing spaces; any model from Status Page) |
| Base API URL | https://llm.ccs.uky.edu/v1 |
| Organization | UKY (or leave blank) |
| Proxy | (leave blank) |
Embedding Model
| Field | Exact Value |
|---|---|
| Embedding model | OpenAI (general interface) :: * (select from dropdown) |
| Local model ID | ccs/Qwen/Qwen3-Embedding-8B (must be an embedding model, NOT a chat model) |
| Base API URL | https://llm.ccs.uky.edu/v1 |
Inline Completions & API Keys
| Field | Exact Value |
|---|---|
| Inline completion model | None (or configure same as language model for code autocomplete) |
| OPENAI_API_KEY | Your key from Locksmith (e.g., sk-abc123...) - personal, or a project key for that project's cloud models. Inside an OOD cluster job, the job's ccs-llm-... key instead. |
Common Errors
AssertionError: model_id was not specified- The Model ID field is empty. Type a model name like
ccs/llama3.1:8b. Model 'ccs/llama3.1:8b ' is not available(note the trailing space)- There is a space after the model name. Delete the trailing space.
Incorrect API key ... platform.openai.com- The request went to OpenAI instead of our gateway. The Base API URL is empty or was not saved. Set it to
https://llm.ccs.uky.edu/v1and click Save Changes again. - Embedding errors with
ccs/llama3.1:8b - You used a chat model for the embedding model. Change it to
ccs/Qwen/Qwen3-Embedding-8B.
Settings are stored per-user in ~/.local/share/jupyter/jupyter_ai/config.json. Each user on the cluster sets their own API key. Model ids are case-sensitive and contain slashes and colons; copy them exactly as GET /v1/models returns them.
Jupyter Notebooks (Python SDK)
# Install: pip install openai
from openai import OpenAI
client = OpenAI(
base_url="https://llm.ccs.uky.edu/v1",
api_key="sk-YOUR_KEY_HERE"
)
# --- List available models ---
models = client.models.list()
for m in models.data:
print(m.id)
# --- Simple completion ---
response = client.chat.completions.create(
model="ccs/llama3.1:8b",
messages=[
{"role": "system", "content": "You are a data science tutor."},
{"role": "user", "content": "Explain PCA in simple terms."}
],
temperature=0.7,
max_tokens=500
)
print(response.choices[0].message.content)
print(f"Tokens used: {response.usage.total_tokens}")
# --- Streaming (good for long responses) ---
stream = client.chat.completions.create(
model="ccs/llama3.1:8b",
messages=[{"role": "user", "content": "Write a Python function to compute Fibonacci numbers."}],
stream=True
)
for chunk in stream:
# The stream ends with a usage-only chunk whose choices list is empty.
content = chunk.choices[0].delta.content if chunk.choices else None
if content:
print(content, end="", flush=True)
# --- Multi-turn conversation ---
conversation = [
{"role": "system", "content": "You are a helpful coding assistant."},
{"role": "user", "content": "What is a decorator in Python?"},
]
r1 = client.chat.completions.create(model="ccs/llama3.1:8b", messages=conversation)
print(r1.choices[0].message.content)
conversation.append(r1.choices[0].message)
conversation.append({"role": "user", "content": "Can you give me a practical example?"})
r2 = client.chat.completions.create(model="ccs/llama3.1:8b", messages=conversation)
print(r2.choices[0].message.content)
VS Code / Continue Extension
Continue is an open-source AI code assistant for VS Code and JetBrains.
- Install the Continue extension from the VS Code marketplace
- Open the Continue config file:
~/.continue/config.json - Add the CCS AI Inference Gateway as a model provider:
{
"models": [
{
"title": "CCS AI Inference Gateway - Qwen3 32B",
"provider": "openai",
"model": "ccs/qwen3:32b",
"apiBase": "https://llm.ccs.uky.edu/v1",
"apiKey": "sk-YOUR_KEY_HERE"
}
],
"tabAutocompleteModel": {
"title": "CCS Autocomplete",
"provider": "openai",
"model": "ccs/llama3.1:8b",
"apiBase": "https://llm.ccs.uky.edu/v1",
"apiKey": "sk-YOUR_KEY_HERE"
}
}
Tip: You can add multiple models. Use a small, fast model for tab
autocomplete and a larger one for chat. Check GET /v1/models or the
Status Page for the current names. There is no
dedicated code-completion model in the catalog today, so a general chat model is the
right choice for both roles.
Cursor
Cursor is an AI-powered code editor that supports custom OpenAI-compatible endpoints.
- Open Cursor Settings (
Cmd+,orCtrl+,) - Go to Models → OpenAI API Key
- Enter your Locksmith key
- Set Override OpenAI Base URL to:
https://llm.ccs.uky.edu/v1 - Add a model name from our Status Page (e.g.,
ccs/llama3.1:8b)
Claude Code / Anthropic SDK
The gateway also speaks the Anthropic Messages API at
https://llm.ccs.uky.edu/v1/messages. This lets
Claude Code and the
Anthropic Python/TypeScript SDKs target the gateway directly with the
same key you use for the OpenAI endpoint. Pick whichever wire format your
client sends natively. This route is served by the upstream inference gateway,
not translated here.
Current scope
Text and vision prompts, streaming and non-streaming, and tool use
(single and parallel, including tools exposed to the model over MCP) are all supported. How well tool use works depends
entirely on the model you choose, not on this portal.
ccs/qwen3:32b and the
thinking ccs/Qwen/Qwen3.6-35B-A3B-FP8 variant are reasonable
starting points for agentic work. Prefer the thinking variant over the -nothink alias for multi-tool loops: without the reasoning phase, -nothink can silently skip a tool call and fabricate the result.
ccs/deepseek-r1:32b does not support tool calling
and returns 400 if you send tools.
Do not send tool_choice: "none" to either
ccs/Qwen/Qwen3.6-35B-A3B-FP8 alias. It is not honoured:
instead of suppressing tool use, the model's raw tool markup is
returned to you inside content with
finish_reason: "stop", which your parser will read as
an ordinary answer. If you do not want a tool called on a given
turn, leave tools out of that request entirely. The
opposite direction works correctly and is the recommended fix for
a skipped call: tool_choice: "required" on turns that
must hit a tool.
Anthropic's signed thinking blocks cannot be reproduced by
non-Claude models, so that request field does not behave as it would against
Anthropic. Reasoning models such as ccs/deepseek-r1:32b emit their
reasoning inline or in a reasoning field instead, depending on the
model and the endpoint.
Claude Code
Set two environment variables before launching claude:
export ANTHROPIC_BASE_URL=https://llm.ccs.uky.edu
export ANTHROPIC_API_KEY=sk-YOUR_KEY_HERE
claude
Claude Code will call ${ANTHROPIC_BASE_URL}/v1/messages and send
your key as x-api-key. Streaming and tool use are both supported,
so the fully agentic flow (Read, Write, Bash, etc.) works. Pick a
tool-capable model (e.g., ccs/qwen3:32b or
ccs/Qwen/Qwen3.6-35B-A3B-FP8, the thinking variant, which is the safer choice for tool-heavy agent loops). Do not use
ccs/deepseek-r1:32b here: it rejects requests that carry
tools.
Anthropic Python SDK
# pip install anthropic
from anthropic import Anthropic
client = Anthropic(
base_url="https://llm.ccs.uky.edu",
api_key="sk-YOUR_KEY_HERE",
)
msg = client.messages.create(
model="ccs/llama3.1:8b",
max_tokens=512,
system="You are a helpful research assistant.",
messages=[
{"role": "user", "content": "What is gradient descent?"}
],
)
print(msg.content[0].text)
print(f"Input tokens: {msg.usage.input_tokens}, Output tokens: {msg.usage.output_tokens}")
curl
curl -X POST https://llm.ccs.uky.edu/v1/messages \
-H "x-api-key: sk-YOUR_KEY_HERE" \
-H "anthropic-version: 2023-06-01" \
-H "Content-Type: application/json" \
-d '{
"model": "ccs/llama3.1:8b",
"max_tokens": 256,
"messages": [
{"role": "user", "content": "What is the capital of Kentucky?"}
]
}'
Key differences from OpenAI /v1/chat/completions
| Aspect | OpenAI route | Anthropic route |
|---|---|---|
| Endpoint | /v1/chat/completions | /v1/messages |
| Auth header | Authorization: Bearer ... | x-api-key: ... (Bearer also accepted) |
| System prompt | Message with role:"system" | Top-level system field |
max_tokens | Optional | Required |
| Stop sequences | stop | stop_sequences |
| Usage | prompt_tokens / completion_tokens | input_tokens / output_tokens |
| Images | image_url data URL | image block with source.type:"base64" |
The model field must match an id returned by
GET /v1/models (also shown on the
Status Page). Anthropic model names like
claude-opus-4-7 are not remapped. Send the CCS model id
(e.g., ccs/llama3.1:8b).
LangChain
# pip install langchain-openai
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
base_url="https://llm.ccs.uky.edu/v1",
api_key="sk-YOUR_KEY_HERE",
model="ccs/llama3.1:8b",
temperature=0.7,
max_tokens=500
)
# Simple invocation
response = llm.invoke("What is retrieval-augmented generation?")
print(response.content)
# With message history
from langchain_core.messages import HumanMessage, SystemMessage
messages = [
SystemMessage(content="You are an expert on natural language processing."),
HumanMessage(content="Compare BERT and GPT architectures.")
]
response = llm.invoke(messages)
print(response.content)
# Streaming
for chunk in llm.stream("Explain attention mechanisms step by step"):
print(chunk.content, end="", flush=True)
LlamaIndex
# pip install llama-index-llms-openai-like
from llama_index.llms.openai_like import OpenAILike
llm = OpenAILike(
api_base="https://llm.ccs.uky.edu/v1",
api_key="sk-YOUR_KEY_HERE",
model="ccs/llama3.1:8b",
is_chat_model=True,
temperature=0.7,
max_tokens=500
)
# Simple completion
response = llm.complete("Explain vector databases in one paragraph.")
print(response.text)
# Chat
from llama_index.core.llms import ChatMessage
messages = [
ChatMessage(role="system", content="You are a helpful assistant."),
ChatMessage(role="user", content="What are embeddings?"),
]
response = llm.chat(messages)
print(response.message.content)
Streaming Responses
Streaming delivers tokens as they are generated, providing a much better user experience for long responses.
Set "stream": true in your request.
Python (OpenAI SDK)
stream = client.chat.completions.create(
model="ccs/llama3.1:8b",
messages=[{"role": "user", "content": "Write a short essay on AI ethics."}],
stream=True
)
for chunk in stream:
# The stream ends with a usage-only chunk whose choices list is empty.
content = chunk.choices[0].delta.content if chunk.choices else None
if content:
print(content, end="", flush=True)
print() # newline at end
Node.js (OpenAI SDK)
const stream = await client.chat.completions.create({
model: "ccs/llama3.1:8b",
messages: [{ role: "user", content: "Write a short essay on AI ethics." }],
stream: true,
});
for await (const chunk of stream) {
const content = chunk.choices[0]?.delta?.content;
if (content) {
process.stdout.write(content);
}
}
console.log(); // newline at end
Node.js (fetch)
const response = await fetch('https://llm.ccs.uky.edu/v1/chat/completions', {
method: 'POST',
headers: {
'Authorization': 'Bearer sk-YOUR_KEY',
'Content-Type': 'application/json'
},
body: JSON.stringify({
model: 'ccs/llama3.1:8b',
messages: [{role: 'user', content: 'Hello!'}],
stream: true
})
});
const reader = response.body.getReader();
const decoder = new TextDecoder();
while (true) {
const {done, value} = await reader.read();
if (done) break;
const chunk = decoder.decode(value);
// Parse SSE lines: "data: {...}\n\n"
for (const line of chunk.split('\n')) {
if (line.startsWith('data: ') && line !== 'data: [DONE]') {
const data = JSON.parse(line.slice(6));
process.stdout.write(data.choices[0]?.delta?.content || '');
}
}
}
Vision Models
Vision models accept images alongside text in the messages array. Images are sent as
base64-encoded data URLs using the standard OpenAI image_url content-part
format. The vision model in the catalog at the time of writing is
ccs/Qwen/Qwen3-VL-8B-Instruct-FP8; models labelled "Vision" on the
Status Page are the current set.
Message Format for Images
Instead of a plain string, set the content field to an array containing text and image parts:
{
"model": "ccs/Qwen/Qwen3-VL-8B-Instruct-FP8",
"messages": [
{
"role": "user",
"content": [
{
"type": "text",
"text": "What is in this image?"
},
{
"type": "image_url",
"image_url": {
"url": "data:image/png;base64,iVBORw0KGgo..."
}
}
]
}
]
}
Python Example
import base64
from openai import OpenAI
client = OpenAI(
base_url="https://llm.ccs.uky.edu/v1",
api_key="sk-YOUR_KEY_HERE"
)
# Read and encode a local image
with open("photo.png", "rb") as f:
image_b64 = base64.b64encode(f.read()).decode("utf-8")
response = client.chat.completions.create(
model="ccs/Qwen/Qwen3-VL-8B-Instruct-FP8",
messages=[
{
"role": "user",
"content": [
{"type": "text", "text": "Describe this image in detail."},
{
"type": "image_url",
"image_url": {
"url": f"data:image/png;base64,{image_b64}"
},
},
],
}
],
max_tokens=500,
)
print(response.choices[0].message.content)
# Output: "The image shows a bar chart with three colored bars representing
# quarterly revenue data. The x-axis labels show Q1, Q2, and Q3..."
Node.js Example
import OpenAI from "openai";
import { readFileSync } from "node:fs";
const client = new OpenAI({
baseURL: "https://llm.ccs.uky.edu/v1",
apiKey: "sk-YOUR_KEY_HERE",
});
// Read and encode a local image
const imageBuffer = readFileSync("photo.png");
const imageB64 = imageBuffer.toString("base64");
const response = await client.chat.completions.create({
model: "ccs/Qwen/Qwen3-VL-8B-Instruct-FP8",
messages: [
{
role: "user",
content: [
{ type: "text", text: "Describe this image in detail." },
{
type: "image_url",
image_url: {
url: `data:image/png;base64,${imageB64}`,
},
},
],
},
],
max_tokens: 500,
});
console.log(response.choices[0].message.content);
// Output: "The image shows a bar chart with three colored bars..."
curl Example
# Encode the image (macOS / Linux)
IMAGE_B64=$(base64 photo.png | tr -d '\n')
curl -X POST https://llm.ccs.uky.edu/v1/chat/completions \
-H "Authorization: Bearer sk-YOUR_KEY_HERE" \
-H "Content-Type: application/json" \
-d '{
"model": "ccs/Qwen/Qwen3-VL-8B-Instruct-FP8",
"messages": [{
"role": "user",
"content": [
{"type": "text", "text": "What is in this image?"},
{"type": "image_url", "image_url": {"url": "data:image/png;base64,'"$IMAGE_B64"'"}}
]
}]
}'
Note: Only models with vision capabilities can process images. Sending images to a text-only model will result in an error. Check the Status Page to see which models support vision.
Model Types
The catalog contains several categories of model, each suited for different tasks.
Understanding the differences helps you choose the right one. The examples below are
illustrative of the catalog as it stands today; it changes, so confirm with
GET /v1/models.
Chat Models
General-purpose language models designed for text generation, conversation, question answering, summarization, and reasoning tasks. These are the most commonly used models and are a good default choice.
Examples: ccs/llama3.1:8b, ccs/qwen3:32b, ccs/gpt-oss:120b
Vision Models
Multimodal models that can accept both text and images as input. Use these when you need a model to analyze, describe, or answer questions about images. Images are sent as base64-encoded data in the messages array (see the Vision Models section above for the exact format).
Example: ccs/Qwen/Qwen3-VL-8B-Instruct-FP8
Reasoning Models
Models that produce an explicit chain of thought before their final answer. This
generally improves accuracy on multi-step problems at the cost of latency and tokens.
Depending on the model and endpoint, the reasoning arrives either inline in the text or
in a separate reasoning field (delta.reasoning when streaming);
clients such as Open WebUI render it in a collapsible section.
Example: ccs/deepseek-r1:32b. Note it does not support tool
calling: sending tools returns 400.
Some models ship a paired -nothink variant, for example
ccs/Qwen/Qwen3.6-35B-A3B-FP8 and
ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink. These are the same weights; the
-nothink id simply has
chat_template_kwargs.enable_thinking = false applied, suppressing the
thinking phase. Prefer -nothink for latency-sensitive and non-tool calls, where the thinking output
is just noise. For agent and tool-calling loops (including MCP-driven agents) prefer the thinking variant instead: in testing, -nothink skipped a tool call and fabricated the answer in about one in five multi-tool turns (the thinking variant: none).
Code Work
There is no dedicated code-completion model in the catalog today. General chat models handle code generation, completion, and debugging, and are what the Continue and Cursor guides above configure. If you need a code-specialised model, ask CCS.
Embedding Models
These models convert text into fixed-length numerical vectors (embeddings). They do not generate text. Instead, the vectors they produce capture the semantic meaning of the input, so similar texts have vectors that are close together in the embedding space. Embeddings are used for:
- Semantic search: find documents similar to a query
- RAG (retrieval-augmented generation): retrieve relevant context before generating an answer
- Clustering and classification: group or categorize text by meaning
Embedding models use the /v1/embeddings endpoint, not /v1/chat/completions. See the Embeddings section for usage details.
Example: ccs/Qwen/Qwen3-Embedding-8B, which requires
encoding_format: "float" on every request.
Send embeddings only to the embedding model. Some chat models will answer /v1/embeddings with a 200 and a correctly-shaped vector that is not a usable embedding. For example, ccs/llama3.1:8b returns a 4096-dimension vector, the same size as the real embedder, so a client cannot tell by shape. Always use ccs/Qwen/Qwen3-Embedding-8B for embeddings.
Context Window Sizes
How much text each model can take in one request differs from model to model, and is covered plainly in Which model should I use?, with a figure in pages for each one. Every model now refuses over-length input with a clear error naming the limit (nothing is sent and nothing is silently dropped) so the only thing to check before you paste is that a model's limit is big enough for your text.
Available Models
Local models (ccs/) run on CCS GPU servers behind the inference gateway;
azure/ models are commercial cloud models reachable on a project key.
The catalog is not fixed: models are added and removed as GPU hosts are re-provisioned, so treat
any list printed in documentation (including this page) as a snapshot.
The authoritative list is GET /v1/models. Query it at
runtime rather than hard-coding ids, and surface the result in your app's model picker
where you can. The Status Page shows the same list
with live up/down state.
Model ids are case-sensitive and contain slashes and colons
(ccs/Qwen/Qwen3-Embedding-8B). Copy them exactly, with no trailing
whitespace.
Model-specific behaviour worth knowing
| Model | What to know |
|---|---|
ccs/deepseek-r1:32b |
Reasoning model. Does not support tool calling: a request carrying tools returns 400. Use a different model for agent workflows. |
ccs/Qwen/Qwen3-Embedding-8B |
Embeddings only. Use /v1/embeddings, and set encoding_format: "float" or the request is rejected. |
ccs/Qwen/Qwen3-VL-8B-Instruct-FP8 |
Vision-capable. Accepts image_url content parts as base64 data URLs. |
ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink |
Identical to ccs/Qwen/Qwen3.6-35B-A3B-FP8 apart from chat_template_kwargs.enable_thinking = false, which suppresses the thinking phase. Preferred for latency-sensitive, non-tool calls. Not recommended for multi-tool agent loops: in testing it skipped a tool call and fabricated the answer in about one in five multi-tool turns (the thinking variant: none). Use ccs/Qwen/Qwen3.6-35B-A3B-FP8 for tool-calling. |
This table covers quirks we have hit in practice; it is not a complete capability matrix. Anything not listed behaves the way the model and the inference gateway behave; this portal adds nothing.
Currently Available
Rendered live from the gateway each time this page loads. This is the catalog as of right now, not a fixed list.
To list models programmatically:
curl https://llm.ccs.uky.edu/v1/models -H "Authorization: Bearer sk-YOUR_KEY_HERE"
Troubleshooting
Errors raised here (auth, key permissions, upstream connectivity) come back as OpenAI-style error JSON. Errors raised by the model or the inference gateway pass through unchanged, so their wording is whatever upstream sent.
401 Unauthorized- "Missing API key" / "Invalid or expired API key"- No key was sent, or the key is wrong, expired, or revoked. Send it as
Authorization: Bearer <key>orx-api-key: <key>. Personal and project keys are managed in Locksmith, which shows whether a key has expired or been revoked; a cluster job key stops working as soon as its SLURM job ends. 403- "key not allowed to access model"- This key cannot use that model. Commercial cloud models through Azure (
azure/...) need your project key; local models (ccs/...) work on your personal key. The message lists the model groups the key you sent can reach. A cluster job key can never reach the cloud models. Switch keys; asking CCS to widen a key is not the fix. 400- rejected by the model- The inference gateway or the model refused the request body. The most common
causes here are sending
toolstoccs/deepseek-r1:32b, which does not support tool calling, and omittingencoding_format: "float"on an embeddings call. The response body carries the upstream message. Read it. 404- unknown model- The model id does not exist in the current catalog. Ids are case-sensitive and
include slashes and colons. Call
GET /v1/modelsfor the live list -- models come and go as GPU hosts are re-provisioned, so an id that worked last month may be gone. 429-rate_limit_error, or the model is busy- Too fast, or too much at once (see Is there a rate limit?),
or every worker is busy generating. Either way the response carries a
Retry-Afterheader, which the official OpenAI and Anthropic SDKs honour automatically; if you wrote your own client, back off for that many seconds and retry. Waiting fixes this one. 429-budget_exceeded- This project's cloud budget is spent. Local models still work on this key. Contact your project PI to add funds.
The body says"Budget has been exceeded! ... Current cost: ..., Max budget: ..."and the response carriesx-should-retry: falseand noRetry-After: waiting does not fix this one, so stop retrying. The same message, without aTeam=name, means your share of the project's pool is used up rather than the whole pool; the PI sets that cap too. Pools are one-off grants unless the PI arranged a periodic reset; the PI sees the balance and any reset date in Locksmith. Only cloud calls are refused -ccs/...models keep answering on the very same key. 502- "Upstream gateway connection failed"- This host could not reach the inference gateway. Check
https://llm.ccs.uky.edu/health/readiness. Ifchecks.upstreamisfalse, the outage is upstream rather than yours. Retry in a few minutes. 504- "Upstream gateway timeout"- No response within the 10-minute read timeout. Reduce
max_tokens, pick a smaller model, or stream. Streaming also stops intermediaries idling the connection out. 503- "Gateway not configured"- The service is missing its upstream credential. That is an operator problem, not yours. Contact CCS.
- Connection refused / cannot reach server
- Check
https://llm.ccs.uky.edu/health/livenessfrom your network - it should return200. This address is reachable from the public internet, so no VPN is required and it works from off campus.
If that succeeds but your request still fails, the problem is the request rather than the network - check the API path and your key.
Known Limitations
This portal does not implement the API. It authenticates you and forwards the request verbatim to the CCS inference gateway, which serves the models. So what works is a property of the model you chose and the inference gateway, not of anything here. We add no samplers, no defaults, and no model filtering, and we do not rewrite responses.
The one exception: on streaming OpenAI-style requests we set
stream_options.include_usage if you did not, so the final SSE chunk carries
real token counts for usage accounting. If you already set it, yours is used.
The table below reflects what we have observed in practice. Treat it as guidance, not a contract. Verify against the model you actually intend to use.
| Feature | Status | Notes |
|---|---|---|
| Chat completions | Works | Streaming and non-streaming |
| Embeddings | Works | Pass encoding_format: "float" explicitly; dimensions is ignored |
| Streaming (SSE) | Works | Proxied with buffering disabled, so chunks arrive as the model produces them |
temperature, top_p, max_tokens, stop | Works | Standard OpenAI sampling fields, honoured by the models in the catalog |
| Vision (images) | Model-dependent | Works with a vision-capable model such as ccs/Qwen/Qwen3-VL-8B-Instruct-FP8. Base64 data URLs; we do not fetch remote image URLs on your behalf. |
| Tool / function calling | Model-dependent | Standard tools / tool_choice. Quality varies sharply by model; the -nothink variants and ccs/qwen3:32b are the practical choices. ccs/deepseek-r1:32b rejects tools with 400. |
| Thinking / reasoning | Model-dependent | Reasoning models expose their chain of thought as a reasoning field, a delta.reasoning when streaming, or inline in the text, depending on the model. Not every client renders it. Use a -nothink variant to switch it off. |
frequency_penalty, presence_penalty, seed, logprobs, response_format | Varies | Forwarded untouched. Whether a given model and serving engine honours, ignores, or rejects each one is decided upstream. Test before relying on it. |
Engine-native sampler params (top_k, min_p, repeat_penalty, ...) | Varies | Not part of the OpenAI spec. Forwarded as sent; acceptance depends entirely on the upstream serving engine. There is no alternative native endpoint exposed through this host. |
| Anthropic Messages API | Works | /v1/messages is served upstream for Claude Code and the Anthropic SDKs: text, vision, streaming and tool use. Anthropic's signed thinking blocks cannot be reproduced by non-Claude models. |
| Model catalog | Changes | Models are added and removed as GPU hosts are re-provisioned. Query GET /v1/models; do not hard-code ids. |
If a parameter does not behave as you expect, the place to look is the model and the inference gateway. Report it to CCS with the model id and the exact request body and we will chase it upstream.
Frequently Asked Questions
- Are my prompts and responses logged?
- Not by this host. We record metadata only - timestamp, model, status code, token counts, latency, client address, and the last four characters of the key; prompt and response content are never stored.
- Where do I get an API key?
- From locksmith.ccs.uky.edu. This portal does not issue personal keys. Keys for OOD cluster jobs are minted automatically by the job and are bound to it.
- How many API keys can I have, and how long do they last?
- Both are set by Locksmith policy. Manage, extend, and revoke your keys there.
- What happens when my key expires?
- Requests return
401 Unauthorized. Renew or reissue in Locksmith. - Can I use this from off-campus?
- Yes.
https://llm.ccs.uky.eduis reachable from the public internet and no VPN is needed. Your API key is what authenticates you, not your network location.
The one exception is cluster job provisioning at/api/v1/*, which stays campus-only onllm-internal.ccs.uky.edu. That is called by SLURM jobs on LCC / MCC / ECC, not by people. - What models are available?
- Call
GET /v1/models, or check the Status Page. The catalog changes as GPU hosts are re-provisioned, so query it rather than hard-coding ids. - Can I request a specific model?
- Contact CCS. Models run on CCS GPU servers, so whether a given model can be added depends on GPU memory, the serving engine, and what else is deployed. It is a capacity conversation, not a download.
- Is there a rate limit?
- Yes, per API key, in two forms: a limit on how many requests you can start
per second, and a limit on how many you can have in flight at once. There is no
cap on total requests per day and no token quota.
Both are set well above what the GPUs can actually serve, so they exist to stop a runaway script, not to pace normal work. In practice you will not reach them.
The exact figures are deliberately not published here: they are tuned to the hardware behind the service and move as capacity is added. What stays constant is the contract -- if you hit a limit you get a429with aRetry-Afterheader, and honouring it is all you need to do. If you have a workload that genuinely needs more than one key allows, contact CCS rather than working around it with extra keys: the service is shared, and the limits are what keep it answering for everyone. - I got a
429. Am I being rate limited? - Read the
typein the body first. If it saysbudget_exceeded, this is not about speed at all: a project's cloud pool is spent - see the budget 429, and do not retry. Otherwise a429means one of two quite different things:- You are sending too fast - starting requests faster, or holding more in flight at once, than one key is allowed. Under your control: slow down, or run fewer in parallel.
- The model is busy - every request is holding a worker while it generates. Not under your control: the fix is to wait, reduce how many requests you run in parallel, or use a model that is already loaded. See the status page for what is running now.
429, it is capacity, and sending them more slowly one at a time will help where sending them all at once will not.
Either way the response carries aRetry-Afterheader. The official OpenAI and Anthropic SDKs honour it automatically; if you wrote your own client, wait that many seconds before retrying. Please be mindful of shared resources. - Can I call the API directly from a web page?
- No, and this is deliberate rather than an oversight. A browser will refuse a
cross-origin request to
/v1from another site.
The reason is that to make such a call, the page has to carry your API key in its JavaScript, where anyone who opens the developer tools can read it. Your key identifies you personally, so publishing it is not a small mistake.
Call the API from your server instead and have your page talk to that. Your key stays on a machine you control, and it is what the official OpenAI and Anthropic SDKs assume you are doing.
The one exception is our own test page, which works because it is served from this site - a page calling back to its own origin never involves cross-origin rules at all. - Who can I contact for help?
- Email the CCS team or visit ccs.uky.edu for support contacts.