OperationalThe upstream inference gateway is reachable.
8 respondingEvery model is probed with one minimal request every 15 minutes;
the last sweep ran 12 min ago. “Responding”
means the model answered that probe (it checks that the model
replies, not that the answer is correct) and it
is a recorded result, not a call made when you loaded this page.
Four NVIDIA GH200 Grace Hopper servers serve these models. Each pairs an H100-class GPU (96 GB HBM3) with a 72-core Grace CPU and ~480 GB of LPDDR5X - about 576 GB of unified memory over NVLink-C2C.
Reasoning models spend tokens thinking before they answer
A reasoning model uses part of your max_tokens budget on internal reasoning. If the budget is small it can spend ALL of it thinking and return an empty string -- which looks like a broken service but is not. Measured here: ccs/qwen3:32b used 289 completion tokens to answer "What is 2+2". At max_tokens=24 the same question returned "". Ask reasoning models for at least 2,000 tokens, or use a non-thinking model such as ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink or ccs/llama3.1:8b.
chat
General instruction-following. Ask it something, get an answer.
reasoning
Thinks step by step before answering. Better at hard problems, slower, and it spends tokens on the thinking.
vision
Accepts images alongside text.
embeddings
Turns text into vectors for search and RAG. Does NOT answer questions.
Qwen3.6 35B-A3B (no-think)
Responding
chat
Same model, thinking switched off. The fastest thing here.
Interactive chat and autocomplete where latency is what you feel
Claude Code and similar tools that dislike a thinking preamble
Short, well-specified tasks that do not need deliberation
Detail, limits and gotchas
Notes
Measured fastest of all eight: about 0.10s to a short non-streaming reply.
Not a different model. If you want deliberation, use the thinking alias instead.
Deployment
Alias of ccs/Qwen/Qwen3.6-35B-A3B-FP8, same weights, same GPU, reasoning disabled.
Limit is about 200 pages (approximately 131,000 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.
High-volume jobs where per-request latency matters
A cheap first pass before escalating to a bigger model
Detail, limits and gotchas
Watch out
This is NOT a search / embedding model, but it will still accept a 'turn this text into search numbers' request and hand back 4,096 numbers that look real and are not. A search tool built on them quietly returns poor results, and nothing in the reply reveals the problem. Send embedding and semantic-search requests only to ccs/Qwen/Qwen3-Embedding-8B.
Notes
No thinking phase, so what you set as max_tokens is what you get as answer.
Deployment
Limit is about 25 pages (approximately 15,500 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.
Agentic and tool-using workflows -- this is what it was built for
Coding and code review
Anything where you want reasoning quality without a 120B latency bill
Detail, limits and gotchas
Notes
Only 3B of the 35B parameters run per token, which is why it is the fastest capable model here.
It thinks before answering. Budget max_tokens accordingly -- see the reasoning-token warning below.
Deployment
256 experts, 8 routed + 1 shared active per token
Limit is about 200 pages (approximately 131,000 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.
NO TOOL CALLING. Sending `tools=` returns HTTP 400. This is empirically verified, not a guess -- its Modelfile has no tools block. If you need tools plus reasoning, use Qwen3.6-35B-A3B-FP8.
This is NOT a search / embedding model, but it will still accept a 'turn this text into search numbers' request and hand back 5,120 numbers that look real and are not. A search tool built on them quietly returns poor results. Send embedding and semantic-search requests only to ccs/Qwen/Qwen3-Embedding-8B.
Notes
Writes long. It will use the whole max_tokens budget on reasoning if you let it.
Deployment
Limit is about 45 pages (approximately 28,000 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.
Supports adjustable reasoning effort upstream; not currently exposed as a parameter here.
A large model that uses a full GPU on its own.
Deployment
Limit is about 190 pages (approximately 124,000 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.
Reasoning tasks where you want a dense model rather than an MoE
Comparing against the MoE models on the same prompt
Detail, limits and gotchas
Notes
Thinks a LOT. Measured: 289 completion tokens to answer 'What is 2+2'. Give it room.
Deployment
Limit is about 25 pages (approximately 15,500 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.
Describing figures, charts, screenshots and scanned pages
Pulling text and structure out of an image
Any prompt that mixes an image with a question
Detail, limits and gotchas
Notes
Pass images as an `image_url` content part, the standard OpenAI vision shape. A base64 data URI works.
The only model here that accepts images. The others will ignore or reject them.
Deployment
Limit is about 100 pages (approximately 65,000 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.
Last check 12 min agoLatency 57msUptime 24h 100.0% (96 checks)
Qwen3 Embedding 8B
Responding
embeddings
Turns text into vectors. It does NOT answer questions.
ccs/Qwen/Qwen3-Embedding-8B
Size8BArchDense embedding modelContextabout 50 pages per item (approximately 33,000 tokens)Dimensions4096Endpoint/v1/embeddings
Good for
RAG -- embed your documents, embed the question, retrieve the closest chunks
Semantic search over a corpus
Clustering, deduplication, similarity scoring
Detail, limits and gotchas
Watch out
Pass `encoding_format: "float"`. It used to be required -- without it the request failed with HTTP 400 -- and since the upstream upgrade on 2026-08-16 it succeeds either way. Keep sending it so your code does not depend on which upstream version happens to be deployed.
Sending a chat request to this model triggers an upstream cooldown: every caller of this model then gets HTTP 429 for roughly 30 seconds, not just the person who made the mistake. The model is not faulty when that happens and it recovers on its own. Measured 2026-08-16.
Not for
Chat or question answering. It has no answer to give -- it returns numbers, not text.
Do not send it to /v1/chat/completions. Use /v1/embeddings. A chat request to this model does not only fail for you -- it takes the model out for everyone for about 30 seconds.
Notes
Returns a 4096-dimension vector per input. Size your vector store for that.
Typical RAG shape: embed chunks once and store them, embed each query at ask time, retrieve top-k by cosine similarity, then pass those chunks to a chat model as context.
Deployment
Limit is about 50 pages per item (approximately 33,000 tokens) of input: every message in the conversation added together, including the system prompt and anything retrieved. The reply you ask for does not count. Go over it and you get a clear error naming the limit: nothing is sent, and nothing is silently dropped.
Last check 12 min agoLatency 75msUptime 24h 100.0% (96 checks)
Using the API
The gateway speaks the OpenAI API. Point any OpenAI-compatible client
at the base URL below, authenticate with a Bearer token, and pass one
of the model ids above as model.
Base URLhttps://llm.ccs.uky.edu/v1
This host issues no keys to people. For a durable personal key, use
Locksmith. If you are running an
Open OnDemand session or a batch job, a short-lived key is issued to
the job automatically and revoked when the job ends. There is
nothing to do.
Never share your key or embed it in published code. To check a key
and compare models without installing anything, use
Try it. See the
full documentation for
streaming, embeddings, and client integrations.