
pytest-mockllm
$ pip install "pytest-mockllm[openai]"The Challenge
Live LLM calls introduce variable responses, latency, and external dependencies into tests. Teams need controlled examples of successful responses, streams, rate limits, and failures.
The Solution
Built fixture-scoped mocks for supported SDK methods, including synchronous and asynchronous responses. Added repeatable failure simulation, call tracking, and token estimates. Recording and replay are currently unavailable. Active fixtures intercept supported paths; installing the plugin does not establish a process-wide network boundary.
Fixture-Scoped LLM Test Harness
Active provider fixtures supply configured responses on supported SDK paths; installing the plugin alone does not block network access.
- ✓Fixture-Scoped Interception
- ✓Supported Sync & Async SDK Paths
- ✓Typed OpenAI & Anthropic Responses
- ✓Streaming Scenarios
- ✓Deterministic Error Simulation
- ✓Token & Cost Estimates
Case Study: Fixture-Scoped LLM Testing with pytest-mockllm
Updated September 5, 2026 to match the current project documentation. This replaces the earlier v0.2.1 case study; existing links remain valid.
The Challenge
Live provider calls introduce variable responses, latency, credentials, and external dependencies into application tests. Developers need repeatable responses and failure scenarios to check parsing, retries, streaming, and tool routing.
The Solution
pytest-mockllm supplies explicit pytest fixtures for supported provider interfaces. Tests configure responses and inspect call details locally. OpenAI and Anthropic supported paths return official SDK response or event objects when those SDKs are installed.
Interception is fixture-scoped
Installing the plugin alone does not prevent live requests. A test must request the appropriate fixture, such as mock_openai, and use a supported interface. Active OpenAI and Anthropic fixtures also reject unhandled SDK base requests. Gemini and LangChain interception covers documented high-level entry points, not every network path.
Keep real credentials out of unit tests. Use separate CI egress controls when the entire suite must remain offline.
Supported behavior
- Configured synchronous and asynchronous responses on documented SDK paths.
- Streaming scenarios and supported tool-call response objects.
- Queued responses, strict mode, and call tracking.
- Deterministic error simulation for repeatable failure cases.
- Token and cost estimates, which are not provider billing measurements.
The Gemini integration targets the legacy google-generativeai SDK. Its replacement, google-genai, is not implemented yet. Consult the current README for the exact supported interfaces.
Recording and Replay Are Unavailable
Current recording and replay modes fail before test code can reach a provider. They do not create or replay cassettes. Earlier versions exposed an interface without safely intercepting provider calls; use deterministic provider fixtures until recording returns with provider-level behavioral tests.
Redaction cannot guarantee that arbitrary data is safe to share. The earlier cassette-safety and zero-leak claims have been withdrawn.
What a Passing Test Establishes
A passing mock test checks application behavior against configured inputs. It does not establish live model quality, complete provider compatibility, process-wide network isolation, or universal absence of flakiness.
Maintain separate, explicit provider contract tests and live evaluations for those different questions.
Try It
python -m pip install "pytest-mockllm[openai]"
See the current README and quick start for a complete fixture-based example and PyPI for published releases.