finstat (local DeepSeek-V4.1-Flash): wrong answers with thinking off
Tested 2026-10-02. Endpoint 192.168.54.242:8000 (SGLang), same results through the :8001 proxy.
Result
Same prompts sent to local finstat and to deepseek/deepseek-v4.1-flash on OpenRouter. Prompt token counts are identical (244 / 527 / 1176 with thinking off), so both servers get the same input. Each cell is correct answers out of 4 runs.
| Setup | 200 tok | 500 tok | 1000 tok |
|---|---|---|---|
| OpenRouter, thinking off | 4/4 | 4/4 | 4/4 |
| OpenRouter, thinking on | 4/4 | 4/4 | 4/4 |
| Local, thinking off (server default) | 1/4 | 0/4 | 0/4 |
Local, chat_template_kwargs: {"thinking": true} | 2/4 | 4/4 | 4/4 |
Local, reasoning_effort: "high" | 3/4 | 4/4 | 4/4 |
- Prompts under ~170 tokens: local works fine.
- 2000-token prompt, thinking off: local finds the code 4/4 but ignores the second question.
- With thinking on, local still fails the 200-token JSON extraction 1-2 times out of 4. OpenRouter gets it right every time.
Example of a failure
500-token prompt: a report with a code (BLUE-HERON-7731) hidden in section 6, asking for the code plus a 2-sentence summary.
- OpenRouter, thinking off:
(1) BLUE-HERON-7731 (2) The backup generator in Barn C had not been tested since March... - Local, thinking off, temperature 0:
I can't help with that request.
Other local failures were refusals, "you didn't ask a question", "the text is garbled", answers in Chinese, and empty responses.
Server settings that differ from a plain setup (/get_server_info)
kv_cache_dtype = fp8_e4m3enable_deepseek_v4_fp4_indexer = truespeculative_algorithm = DSPARK(6 draft tokens, draft model = the main model)- SGLang version
0.0.0.dev1+g6407cc519
None of these is confirmed as the cause.
Prompts used: finstat-test-prompts.md.