Skip to main content

LLM Testing in Python with Pytest

How to test LLM-powered Python applications with pytest using deterministic mocks, datasets, evaluators, thresholds, async tests, and redaction.

3 min read
LLM Testing in Python with Pytest
On this page

Testing LLM-powered Python code with live model calls in every test is a trap. The suite becomes slow, expensive, flaky, and unsafe to run in CI. The answer is not to avoid testing. The answer is to separate deterministic system tests from smaller live evaluation runs.

Langfuse frames the practical architecture as datasets, experiment runners, and evaluators. That maps cleanly to pytest if you treat model behavior as something to score or mock, not as a string to compare blindly.

What to unit test

Unit tests should cover deterministic code around the model.

  • Prompt construction.
  • Tool routing.
  • JSON/schema parsing.
  • Retry and timeout behavior.
  • PII redaction.
  • Cost and token-budget guards.
  • Refusal/error handling.
  • Fallback provider selection.

Mock the LLM response for these. If the parser breaks on malformed JSON, you do not need a real model to discover it.

What to evaluate

Evaluation tests check behavior across a dataset. A test case has input, expected behavior, and an evaluator. The evaluator can be code-based, semantic, or LLM-as-judge. For objective tasks, use code. If the expected answer is “Paris,” do not use a judge model when a case-insensitive contains check is enough.

The following integration sketch assumes your application supplies SupportBot and a fake client that records the messages actually sent. Adapt those interfaces to your code; the test must inspect the outgoing payload, not merely the displayed response.

python
def test_redaction_happens_before_llm_call(fake_client):
    app = SupportBot(client=fake_client)
    app.answer("My card is 4111-1111-1111-1111")
    sent_prompt = fake_client.last_messages[-1]["content"]
    assert "4111-1111-1111-1111" not in sent_prompt
    assert "[REDACTED_CARD]" in sent_prompt

Thresholds are normal

Choose evaluation thresholds from the consequence of a failure and a labeled validation set. Critical authorization and data-integrity rules may require zero tolerated violations; a stylistic preference can use a different bar. Report sample size and failure categories alongside the aggregate score. There is no universal 95% or 80% threshold for LLM applications.

Keep live tests explicit

Live provider tests should be opt-in, labeled, and budgeted. Run them nightly or before release, not on every file save. Store representative datasets. Track results over time. If a model upgrade improves one behavior and breaks another, you need historical comparison, not vibes.

Why pytest still matters

Pytest is excellent glue. Fixtures isolate clients. Parametrization runs datasets. Markers separate live from mocked tests. CI turns failures into gates. The LLM part is new; the engineering discipline is not.

Sources and further reading

Dhiraj Das

About the Author

Dhiraj Das is an Automation Consultant with over a decade of experience building systems that expose failures, reduce flakiness, and make complex workflows repeatable. He applies that discipline to AI-agent validation, LLM testing, and postmortems.

He shares small open source utilities from real automation work, including: waitless (flaky tests), sb-stealth-wrapper (bot detection), selenium-teleport (state persistence), selenium-chatbot-test (AI chatbot testing), lumos-shadowdom (Shadow DOM), and visual-guard (visual regression).

Share this article: