Ran a second live smoke test — load_model() then chat_structured()
with a real pydantic.BaseModel schema — against the same router-mode
llama-server used for the earlier fix. Confirmed pydantic-ai's
Agent/OpenAIProvider mechanism works against llama-server's
OpenAI-compatible endpoint, not just Ollama's (the only one verified
live in the prerequisite epic). No gap found; every LlamaCppProvider
method has now been exercised against a real server, not just mocks.
Also strengthened the two idempotency regression tests added in the
previous commit: they previously asserted only "does not raise", which
would also pass if the method silently no-op'd for an unrelated bug.
Now assert the mocked _post_json was actually called with the expected
request, so the test verifies real behavior, not just absence of a
crash.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_0132ojafeazQ3ephcBejEWFj