fix(litellm): default max_tokens=4096 for the reasoning model
Qwen3 spends output tokens on reasoning_content before writing content. Open WebUI's default chat request doesn't set max_tokens, so it fell through to llama.cpp's low default and the model ran out mid-thought, returning finish_reason=length with empty content — no reply shown in Open WebUI. Confirmed via a manual /v1/chat/completions call: works with max_tokens=2000, fails without it. litellm_params.max_tokens is a default, not a cap — any caller (or Open WebUI's per-model Advanced Params) that sets its own max_tokens still overrides it. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Q2CR8yawSf7pwAVYnjwFea
This commit is contained in:
@@ -7,6 +7,13 @@ model_list:
|
|||||||
model: openai/qwen3.8-27b-local
|
model: openai/qwen3.8-27b-local
|
||||||
api_base: http://llama-server:8080/v1
|
api_base: http://llama-server:8080/v1
|
||||||
api_key: local
|
api_key: local
|
||||||
|
# Qwen3 is a reasoning model — it spends output tokens on
|
||||||
|
# reasoning_content before ever writing content. Callers that don't
|
||||||
|
# set their own max_tokens (Open WebUI's default request didn't) hit
|
||||||
|
# llama.cpp's low default, so the model runs out mid-thought and
|
||||||
|
# content comes back empty. This is a floor, not a cap — any caller
|
||||||
|
# that passes its own max_tokens still overrides it.
|
||||||
|
max_tokens: 4096
|
||||||
model_info:
|
model_info:
|
||||||
# Shadow cloud-cost estimate — priced against Claude Sonnet 5's published
|
# Shadow cloud-cost estimate — priced against Claude Sonnet 5's published
|
||||||
# rate, not real spend (this proxy only ever routes to the local model).
|
# rate, not real spend (this proxy only ever routes to the local model).
|
||||||
|
|||||||
Reference in New Issue
Block a user