Try it
Runs in your browser against the live API. Nothing to install.
Reasoning models spend tokens thinking before they answer
A reasoning model uses part of your max_tokens budget on internal reasoning. If the budget is small it can spend ALL of it thinking and return an empty string -- which looks like a broken service but is not. Measured here: ccs/qwen3:32b used 289 completion tokens to answer "What is 2+2". At max_tokens=24 the same question returned "". Ask reasoning models for at least 2,000 tokens, or use a non-thinking model such as ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink or ccs/llama3.1:8b.
Run the same thing from a terminal
These update as you change the test, key, model and prompt above.curl https://llm.ccs.uky.edu/v1/chat/completions \
-H "Authorization: Bearer YOUR_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink",
"messages": [
{
"role": "user",
"content": "What is 2+2? Answer in one short sentence."
}
],
"max_tokens": 512
}'
pip install openai
from openai import OpenAI
client = OpenAI(
base_url="https://llm.ccs.uky.edu/v1",
api_key="YOUR_KEY",
)
resp = client.chat.completions.create(
model="ccs/Qwen/Qwen3.6-35B-A3B-FP8-nothink",
messages=[{"role": "user", "content": "What is 2+2? Answer in one short sentence."}],
max_tokens=512,
)
print(resp.choices[0].message.content)
# Two variables and every OpenAI-compatible tool works: the openai
# CLI, aider, llm, LangChain, LlamaIndex, Continue, Open WebUI.
# Put them in a job script, not in a file you commit.
export OPENAI_BASE_URL="https://llm.ccs.uky.edu/v1"
export OPENAI_API_KEY="YOUR_KEY"
# Check the key works and see what is running right now:
curl -s "$OPENAI_BASE_URL/models" \
-H "Authorization: Bearer $OPENAI_API_KEY"
A copied line goes into your shell history. Prefix it with a space,
or use the Environment tab, if that matters to you. Full base URL:
https://llm.ccs.uky.edu/v1
· no key yet? Get one from
Locksmith
· more clients in the documentation
· what each model is for, on the
status page.