# finstat (local DeepSeek-V4.1-Flash): wrong answers with thinking off

Tested 2026-10-02. Endpoint `192.168.54.242:8000` (SGLang), same results through the `:8001` proxy.

## Result

Same prompts sent to local `finstat` and to `deepseek/deepseek-v4.1-flash` on OpenRouter. Prompt token counts are identical (244 / 527 / 1176 with thinking off), so both servers get the same input. Each cell is correct answers out of 4 runs.

| Setup | 200 tok | 500 tok | 1000 tok |
|---|---|---|---|
| OpenRouter, thinking off | 4/4 | 4/4 | 4/4 |
| OpenRouter, thinking on | 4/4 | 4/4 | 4/4 |
| Local, thinking off (server default) | 1/4 | 0/4 | 0/4 |
| Local, `chat_template_kwargs: {"thinking": true}` | 2/4 | 4/4 | 4/4 |
| Local, `reasoning_effort: "high"` | 3/4 | 4/4 | 4/4 |

- Prompts under ~170 tokens: local works fine.
- 2000-token prompt, thinking off: local finds the code 4/4 but ignores the second question.
- With thinking on, local still fails the 200-token JSON extraction 1-2 times out of 4. OpenRouter gets it right every time.

## Example of a failure

500-token prompt: a report with a code (`BLUE-HERON-7731`) hidden in section 6, asking for the code plus a 2-sentence summary.

- OpenRouter, thinking off: `(1) BLUE-HERON-7731 (2) The backup generator in Barn C had not been tested since March...`
- Local, thinking off, temperature 0: `I can't help with that request.`

Other local failures were refusals, "you didn't ask a question", "the text is garbled", answers in Chinese, and empty responses.

## Server settings that differ from a plain setup (`/get_server_info`)

- `kv_cache_dtype = fp8_e4m3`
- `enable_deepseek_v4_fp4_indexer = true`
- `speculative_algorithm = DSPARK` (6 draft tokens, draft model = the main model)
- SGLang version `0.0.0.dev1+g6407cc519`

None of these is confirmed as the cause.

Prompts used: `finstat-test-prompts.md`.
