Step 07 - Testing, Evaluation, and Observability
Checking that the planner is right, not just plausible
The trip planner now fetches real weather and points of interest before any language model runs. But the plan it returns is still only as good as the model’s judgment of that data, and a model that produces a convincing itinerary can still miss the point. The plan for a three-day family trip might skip a rest day, the cost line might come out of a miscalculation, and the vehicle recommendation might ignore the budget because nothing in the pipeline checks it.
This step adds a repeatable way to answer three questions that the demo UI cannot:
- Is a given plan structurally sound? A deterministic check catches a missing vehicle, a gap in the itinerary, or an invalid cost without calling a model.
- Does the plan match what the trip actually requires? A judge model compares the output against an expected plan, and a semantic similarity score measures how close they are even when the wording differs.
- Where did the plan come from? The planning run is traced with OpenTelemetry, the trace is exported to Langfuse, and the evaluation score is attached to the exact trace it was measured from.
The result is an evaluation harness you can run before every change to a prompt or a model: the same samples, the same checks, a report with a score and per-case details. When the score drops, the report points at which sample failed and why.
How evaluation differs from the guardrails in Step 03
Step 03 added guardrails that block a request or a response when it violates a hard rule, such as a destination with a known safety issue. Those are enforcement mechanisms that sit in the request path.
Evaluation sits outside the request path. It runs against saved outputs of completed planning runs, it produces a score and a report rather than a decision, and it uses a judge model that is separate from the model that produced the plan. A guardrail stops a bad plan from being returned; evaluation tells you, after the fact, how often the planner produces bad plans and which ones.
flowchart LR
subgraph runtime["Request path (Step 03)"]
req["Plan request"] --> gr["Guardrails<br/><small>block or pass</small>"]
gr --> plan["Trip plan"]
end
subgraph eval["Evaluation path (this step)"]
saved["Saved plan output"] --> inv["Invariant checks<br/><small>deterministic</small>"]
saved --> judge["AI judge<br/><small>separate model</small>"]
inv --> rep["Report + score"]
judge --> rep
plan -.->|"saved once"| saved
end
plan -.->|"trace"| lf["Langfuse<br/><small>score attached</small>"]
Choosing a starting point
Apply the changes below to your Step 06 working project. Keep your existing model-provider settings. Adding quarkus-langfuse and quarkus-opentelemetry to pom.xml triggers an automatic restart in dev mode; the first restart after adding those dependencies will take longer than usual because Dev Services pulls and starts Langfuse containers.
Copy section-3/step-07 to a working directory and open that copy. Apply your model-provider settings. The code changes below are already included, including BudgetVerdict.java, VehicleBudgetJudge.java, the updated TripAppropriatenessGuardrail.java, and the judgeModel configuration. Join the hands-on route at Adding the evaluation dependencies.
The same prerequisites from Step 06 apply: a model provider key, a container runtime for the Dev Services the planner needs, and the Trip Intelligence MCP server running on port 8085 for the live evaluation runs.
Upgrading the vehicle guardrail to an LLM judge
The TripAppropriatenessGuardrail in Steps 02–06 uses a hardcoded list of brand names to detect luxury vehicles on economy budgets. That list covers obvious supercars but misses premium SUVs and saloons that are equally unaffordable on an economy budget. In this step we replace that heuristic with an LLM-based judge that evaluates the vehicle’s type, model name, and reasoning text together, so the decision reflects intent rather than brand-name matching.
Create src/main/java/com/tripplanner/guardrails/BudgetVerdict.java:
package com.tripplanner.guardrails;
/**
* Structured verdict returned by the LLM-based budget guardrail judge.
*
* @param appropriate {@code true} if the vehicle is a reasonable choice for the given budget tier
* @param reason a short explanation of the verdict, used for audit logging and UI display
*/
public record BudgetVerdict(boolean appropriate, String reason) {}
Create src/main/java/com/tripplanner/guardrails/VehicleBudgetJudge.java:
package com.tripplanner.guardrails;
import dev.langchain4j.service.SystemMessage;
import dev.langchain4j.service.UserMessage;
import io.quarkiverse.langchain4j.RegisterAiService;
import jakarta.enterprise.context.ApplicationScoped;
/**
* LLM-as-judge that evaluates whether a vehicle recommendation is appropriate
* for the given budget tier. Used by {@link TripAppropriatenessGuardrail} in
* place of a hardcoded brand allowlist.
*
* <p>Runs on a fast, cheap model (configured as {@code judgeModel}) to keep
* latency low. Returns a structured {@link BudgetVerdict} with a boolean verdict
* and a short explanation suitable for audit logging.
*/
@ApplicationScoped
@RegisterAiService(modelName = "judgeModel")
public interface VehicleBudgetJudge {
@SystemMessage("""
You are a vehicle budget compliance judge.
Your only job is to decide whether a vehicle recommendation is appropriate for a given budget tier.
Economy budgets (roughly €500–€1000 total or €30–€50/day) require affordable, mass-market vehicles.
Premium or luxury vehicles — regardless of brand — are not appropriate for economy budgets.
Be strict: if there is any doubt, mark the vehicle as not appropriate.
Respond only with a JSON object matching the BudgetVerdict schema.
""")
@UserMessage("""
Budget tier: {budget}
Vehicle type: {vehicleType}
Vehicle model: {vehicleModel}
Vehicle reasoning: {reasoning}
Is this vehicle appropriate for the stated budget tier?
""")
BudgetVerdict evaluate(String budget, String vehicleType, String vehicleModel, String reasoning);
}
VehicleBudgetJudge is a @RegisterAiService bound to a named model called judgeModel, and is annotated @ApplicationScoped so it can be injected and called from threads that have no active HTTP request context, such as the Quarkus Flow executor threads that run the guardrail. Without that annotation, the default @RequestScoped CDI proxy would throw a RequestScoped context was not active error every time the guardrail tried to reach it. Its system prompt explains what the judge must decide and instructs it to be strict: if there is any doubt, the vehicle is not appropriate. The user message gives the judge the budget tier, vehicle type, vehicle model, and the agent’s own reasoning, so it can read the full context rather than pattern-matching a single field.
Update src/main/java/com/tripplanner/guardrails/TripAppropriatenessGuardrail.java to inject and use the judge:
package com.tripplanner.guardrails;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.fasterxml.jackson.databind.node.ObjectNode;
import dev.langchain4j.data.message.AiMessage;
import dev.langchain4j.guardrail.OutputGuardrail;
import dev.langchain4j.guardrail.OutputGuardrailRequest;
import dev.langchain4j.guardrail.OutputGuardrailResult;
import jakarta.enterprise.context.ApplicationScoped;
import jakarta.inject.Inject;
import java.util.Set;
import java.util.Locale;
@ApplicationScoped
public class TripAppropriatenessGuardrail implements OutputGuardrail {
private static final Set<String> SMALL_VEHICLE_KEYWORDS = Set.of(
"sports car", "sport car", "coupé", "coupe", "convertible", "2-seater", "two-seater", "roadster");
// Fast pre-check for obvious supercars — avoids a judge call for clear-cut cases.
private static final Set<String> OBVIOUS_LUXURY_BRANDS = Set.of(
"ferrari", "porsche", "lamborghini", "maserati", "bentley", "rolls-royce", "aston martin", "mclaren");
// Brands detected by the user preferences ANNOTATE path (unchanged from earlier steps).
private static final Set<String> LUXURY_BRANDS = Set.of(
"ferrari", "porsche", "lamborghini", "maserati", "bentley", "rolls-royce", "aston martin", "mclaren",
"land rover", "range rover", "jaguar", "bmw", "mercedes", "audi", "lexus");
@Inject
GuardrailAuditLog auditLog;
@Inject
ObjectMapper objectMapper;
@Inject
VehicleBudgetJudge budgetJudge;
@Override
public OutputGuardrailResult validate(OutputGuardrailRequest guardrailRequest) {
String text = guardrailRequest.responseFromLLM().aiMessage().text();
if (text == null || text.isBlank()) {
auditLog.log("TripAppropriatenessGuardrail", "SKIP", "No text content; tool-call content was not validated");
return success();
}
JsonNode root;
try {
root = objectMapper.readTree(TripSafetyGuardrail.extractJson(text));
} catch (Exception e) {
auditLog.log("TripAppropriatenessGuardrail", "SKIP", "Response is not valid JSON; suitability was not validated");
return success();
}
if (root == null || !root.isObject() || !root.path("type").isTextual()
|| root.path("type").asText().isBlank() || !root.path("model").isTextual()
|| root.path("model").asText().isBlank()) {
auditLog.log("TripAppropriatenessGuardrail", "SKIP", "Vehicle type or model is missing; suitability was not validated");
return success();
}
var variables = guardrailRequest.requestParams().variables();
String budget = (String) variables.get("budget");
String tripType = (String) variables.get("tripType");
String preferences = variables.get("preferences") instanceof String s ? s : "";
int travelers;
try {
// Prompt variables may contain text or numeric method arguments.
travelers = Integer.parseInt(String.valueOf(variables.get("travelers")));
} catch (NumberFormatException e) {
auditLog.log("TripAppropriatenessGuardrail", "SKIP", "Traveler count is missing or invalid; suitability was not validated");
return success();
}
if (budget == null || budget.isBlank() || tripType == null || tripType.isBlank()) {
auditLog.log("TripAppropriatenessGuardrail", "SKIP", "Trip variables are missing or incomplete; suitability was not validated");
return success();
}
String vehicleType = root.path("type").asText("").toLowerCase(Locale.ROOT);
String vehicleModel = root.path("model").asText("").toLowerCase(Locale.ROOT);
String reasoning = root.path("reasoning").asText("");
// For economy budgets: fast-path for obvious supercars, then LLM judge for everything else.
if (budget.toLowerCase(Locale.ROOT).contains("economy")) {
boolean inappropriate;
if (isObviousLuxuryBrand(vehicleModel)) {
inappropriate = true;
} else {
try {
BudgetVerdict verdict = budgetJudge.evaluate(budget, vehicleType, vehicleModel, reasoning);
inappropriate = !verdict.appropriate();
auditLog.log("TripAppropriatenessGuardrail",
inappropriate ? "JUDGE-FAIL" : "JUDGE-PASS",
"Budget judge verdict for '" + vehicleModel + "': " + verdict.reason());
} catch (Exception e) {
auditLog.log("TripAppropriatenessGuardrail", "JUDGE-ERROR",
"Budget judge threw an exception for '" + vehicleModel + "' — treating as appropriate: " + e.getMessage());
inappropriate = false;
}
}
if (inappropriate) {
auditLog.log("TripAppropriatenessGuardrail", "REPROMPT",
"Vehicle '" + vehicleModel + "' does not match economy budget — asking model to retry with an affordable option");
return reprompt("The vehicle recommendation does not fit the economy budget. "
+ "Please recommend an affordable, budget-friendly vehicle instead.",
"The previous vehicle recommendation does not fit the economy budget. "
+ "Respond ONLY with a valid JSON vehicle recommendation object (with fields: type, model, reasoning) "
+ "for an affordable, budget-friendly, non-luxury vehicle suitable for this trip. Do not explain or acknowledge — just output the JSON.");
}
}
if (travelers >= 4 && isSmallVehicle(vehicleType)) {
String reason = "Original recommendation '" + vehicleType + "' is too small for " + travelers + " travelers";
rewriteVehicle((ObjectNode) root, travelers, tripType, reason);
auditLog.log("TripAppropriatenessGuardrail", "REWRITE", reason + "; returned a generic category recommendation");
return successWith(AiMessage.from(root.toString()));
}
// If the user asked for a luxury brand but the output doesn't contain it, the model
// already overrode the request on its own. Annotate the result so the UI can show why.
String requestedBrand = requestedLuxuryBrand(preferences);
if (requestedBrand != null && !vehicleModel.contains(requestedBrand)) {
String reason = "Requested brand '" + requestedBrand + "' is not suitable for this trip; a more appropriate vehicle was selected";
((ObjectNode) root).put("guardrailOverride", reason);
auditLog.log("TripAppropriatenessGuardrail", "ANNOTATE",
"Preferences mentioned '" + requestedBrand + "' but output is '" + vehicleModel + "' — " + reason);
return successWith(AiMessage.from(root.toString()));
}
auditLog.log("TripAppropriatenessGuardrail", "PASS", "Budget judge approved the vehicle recommendation");
return success();
}
private boolean isSmallVehicle(String vehicleType) {
return SMALL_VEHICLE_KEYWORDS.stream().anyMatch(vehicleType::contains);
}
private boolean isObviousLuxuryBrand(String vehicleModel) {
return OBVIOUS_LUXURY_BRANDS.stream().anyMatch(vehicleModel::contains);
}
private String requestedLuxuryBrand(String preferences) {
if (preferences == null || preferences.isBlank()) return null;
String lower = preferences.toLowerCase(Locale.ROOT);
return LUXURY_BRANDS.stream().filter(lower::contains).findFirst().orElse(null);
}
private void rewriteVehicle(ObjectNode vehicle, int travelers, String tripType, String reason) {
String replacement = switch (tripType.toLowerCase(Locale.ROOT)) {
case "adventure" -> "SUV";
case "business" -> "Estate";
default -> "MPV";
};
vehicle.put("type", replacement);
vehicle.put("model", (replacement.equals("MPV") ? "Family MPV" : replacement)
+ "; specific model subject to availability.");
vehicle.put("reasoning", "Vehicle corrected by guardrail: original recommendation was too small for "
+ travelers + " travelers. Suggested category: " + replacement
+ ". Confirm seating, luggage capacity, price, and availability with the rental provider.");
vehicle.put("guardrailOverride", reason);
}
}
For economy budgets, the guardrail runs a fast pre-check against OBVIOUS_LUXURY_BRANDS — Ferrari, Lamborghini, and similar supercars — to skip the judge call for clear-cut cases. For everything else it calls the judge and acts on the verdict. This keeps obvious cases cheap while using the full LLM reasoning for the ambiguous ones, such as a Land Rover Discovery recommended for a five-person economy trip.
Add the judgeModel configuration to src/main/resources/application.properties:
# Judge model for LLM-based guardrail evaluation (fast, cheap)
quarkus.langchain4j.judgeModel.chat-model.provider=openai
quarkus.langchain4j.openai.judgeModel.api-key=${OPENAI_API_KEY}
quarkus.langchain4j.openai.judgeModel.chat-model.model-name=gpt-4o-mini
quarkus.langchain4j.openai.judgeModel.chat-model.temperature=0
quarkus.langchain4j.openai.judgeModel.timeout=30
gpt-4o-mini at temperature 0 is fast and deterministic enough for a binary budget verdict. Using a named model keeps its cost separate from the planner’s main model usage, which is visible as a distinct service in Langfuse.
Adding the evaluation dependencies
The evaluation harness uses the quarkus-langchain4j-testing-evaluation modules for sample loading, scoring, and the AI judge. We also need OpenTelemetry for tracing and the Langfuse extension for publishing scores to the trace.
Add the following dependencies to your trip-planner/pom.xml:
<dependency>
<groupId>io.quarkus</groupId>
<artifactId>quarkus-opentelemetry</artifactId>
</dependency>
<dependency>
<groupId>io.quarkiverse.langfuse</groupId>
<artifactId>quarkus-langfuse</artifactId>
<version>${quarkus-langfuse.version}</version>
</dependency>
<dependency>
<groupId>io.quarkus</groupId>
<artifactId>quarkus-junit</artifactId>
<scope>test</scope>
</dependency>
<dependency>
<groupId>io.smallrye.reactive</groupId>
<artifactId>smallrye-reactive-messaging-in-memory</artifactId>
<scope>test</scope>
</dependency>
<dependency>
<groupId>io.quarkiverse.langchain4j</groupId>
<artifactId>quarkus-langchain4j-testing-evaluation-junit5</artifactId>
<scope>test</scope>
</dependency>
<dependency>
<groupId>io.quarkiverse.langchain4j</groupId>
<artifactId>quarkus-langchain4j-testing-evaluation-ai-judge</artifactId>
<scope>test</scope>
</dependency>
<dependency>
<groupId>io.quarkiverse.langchain4j</groupId>
<artifactId>quarkus-langchain4j-testing-evaluation-semantic-similarity</artifactId>
<scope>test</scope>
</dependency>
<dependency>
<groupId>org.awaitility</groupId>
<artifactId>awaitility</artifactId>
<scope>test</scope>
</dependency>
The snippet shows all evaluation-related dependencies. If you are continuing from Step 06, quarkus-junit and smallrye-reactive-messaging-in-memory are already present — add only the ones that are new: quarkus-opentelemetry, quarkus-langfuse, the three testing-evaluation modules, and awaitility.
Saving a plan as text
Everything in this step evaluates one thing: the text form of a completed plan. Before we can check whether a plan is valid or compare it against an expected output, we need a stable way to turn a TripPlan into a string. The invariant checks, the judge comparison, and the test fixtures all operate on this rendering.
Create src/test/java/com/tripplanner/evaluation/TripPlanText.java:
package com.tripplanner.evaluation;
import com.fasterxml.jackson.core.JsonProcessingException;
import com.fasterxml.jackson.databind.ObjectMapper;
import com.tripplanner.model.TripPlan;
public final class TripPlanText {
private static final ObjectMapper JSON = new ObjectMapper();
private TripPlanText() {
}
public static String render(TripPlan plan) {
try {
return JSON.writerWithDefaultPrettyPrinter().writeValueAsString(plan);
} catch (JsonProcessingException e) {
throw new IllegalArgumentException("Cannot save trip plan", e);
}
}
public static TripPlan read(String output) throws JsonProcessingException {
return JSON.readValue(output, TripPlan.class);
}
}
render() serializes a TripPlan to pretty-printed JSON, and read() parses it back. Rendering is deliberately mechanical, so a plan that is structurally broken in the TripPlan shows up as a broken rendering. This is exactly what the checks below look for.
Writing the invariant strategy
The cheapest way to catch a broken plan is to check its shape with plain Java. The invariant strategy inspects a saved plan for the properties every valid plan must have, and it does not call a model at any point, so it runs in milliseconds and costs nothing.
Create src/test/java/com/tripplanner/evaluation/TripPlanInvariantStrategy.java:
package com.tripplanner.evaluation;
import com.tripplanner.model.TripPlan;
import io.quarkiverse.langchain4j.testing.evaluation.EvaluationResult;
import io.quarkiverse.langchain4j.testing.evaluation.EvaluationSample;
import io.quarkiverse.langchain4j.testing.evaluation.EvaluationStrategy;
import java.util.ArrayList;
import java.util.HashSet;
import java.util.List;
public class TripPlanInvariantStrategy implements EvaluationStrategy<String> {
@Override
public EvaluationResult evaluate(EvaluationSample<String> sample, String output) {
if (output == null || output.isBlank()) {
return EvaluationResult.failed(0, "planner returned no output");
}
try {
return evaluate(sample, TripPlanText.read(output));
} catch (com.fasterxml.jackson.core.JsonProcessingException e) {
return EvaluationResult.failed(0, "planner returned invalid plan JSON");
}
}
public EvaluationResult evaluate(EvaluationSample<String> sample, TripPlan plan) {
if (plan == null) return EvaluationResult.failed(0, "planner returned no plan");
int requestedDays = Integer.parseInt(sample.parameters().get(2).toString());
if (requestedDays < 1 || requestedDays > 30) {
throw new IllegalArgumentException("Sample duration must be between 1 and 30 days");
}
List<String> errors = new ArrayList<>();
if (plan.vehicle() == null) {
errors.add("missing vehicle");
} else {
requireText(plan.vehicle().type(), "vehicle type", errors);
requireText(plan.vehicle().model(), "vehicle model", errors);
requireText(plan.vehicle().reasoning(), "vehicle reasoning", errors);
}
requireText(plan.routeOverview(), "route", errors);
if (plan.itinerary() == null || plan.itinerary().size() != requestedDays) {
errors.add("itinerary must contain " + requestedDays + " days");
}
var numbers = new HashSet<Integer>();
if (plan.itinerary() != null) {
for (var day : plan.itinerary()) {
if (day == null) {
errors.add("missing itinerary entry");
continue;
}
if (!numbers.add(day.day())) errors.add("duplicate day " + day.day());
if (day.day() < 1 || day.day() > requestedDays) errors.add("invalid day " + day.day());
requireText(day.title(), "day title", errors);
requireText(day.description(), "day description", errors);
requireText(day.overnightStop(), "overnight stop", errors);
}
}
for (int day = 1; day <= requestedDays; day++) {
if (!numbers.contains(day)) errors.add("itinerary missing day " + day);
}
if (plan.costs() == null) {
errors.add("missing costs");
} else {
String total = plan.costs().total();
if (total == null || !total.matches("\\s*(?:€|EUR)?\\s*\\d+(?:[.,]\\d+)*(?:\\s*(?:€|EUR))?\\s*")) {
errors.add("cost total must be a non-negative EUR amount");
}
}
return errors.isEmpty() ? EvaluationResult.passed(1)
: EvaluationResult.failed(0, String.join("; ", errors));
}
private static void requireText(String value, String field, List<String> errors) {
if (value == null || value.isBlank() || value.equalsIgnoreCase("null")) errors.add("missing " + field);
}
}
The strategy implements EvaluationStrategy<String> from the quarkus-langchain4j-testing-evaluation module. It takes the sample and the actual output, then returns an EvaluationResult with a score from 0 to 1, a pass or fail decision, and a reason. The checks cover the vehicle recommendation fields, the itinerary day count and numbering, and the cost total format.
Pinning the invariants with known-bad fixtures
The invariants need their own test cases so you can verify they catch real defects. Each fixture is a plan with a specific problem, and the test asserts that the strategy fails it and names the defect in the result.
Create src/test/resources/evaluation/known-bad.yaml:
[
{
"name": "missing-vehicle",
"output": "{\"vehicle\": null, \"routeOverview\": \"Rome\", \"itinerary\": [{\"day\": 1, \"title\": \"Day 1\", \"description\": \"Explore\", \"overnightStop\": \"Hotel\"}, {\"day\": 2, \"title\": \"Day 2\", \"description\": \"Explore\", \"overnightStop\": \"Hotel\"}, {\"day\": 3, \"title\": \"Day 3\", \"description\": \"Explore\", \"overnightStop\": \"Hotel\"}], \"costs\": {\"total\": \"\\u20ac310\"}}",
"reason": "missing vehicle"
},
{
"name": "gap-in-days",
"output": "{\"vehicle\": {\"type\": \"MPV\", \"model\": \"Family MPV\", \"reasoning\": \"Room for luggage\"}, \"routeOverview\": \"Rome\", \"itinerary\": [{\"day\": 1, \"title\": \"Day 1\", \"description\": \"Explore\", \"overnightStop\": \"Hotel\"}, {\"day\": 3, \"title\": \"Day 3\", \"description\": \"Explore\", \"overnightStop\": \"Hotel\"}], \"costs\": {\"total\": \"\\u20ac310\"}}",
"reason": "missing day 2"
},
{
"name": "negative-cost",
"output": "{\"vehicle\": {\"type\": \"MPV\", \"model\": \"Family MPV\", \"reasoning\": \"Room for luggage\"}, \"routeOverview\": \"Rome\", \"itinerary\": [{\"day\": 1, \"title\": \"Day 1\", \"description\": \"Explore\", \"overnightStop\": \"Hotel\"}, {\"day\": 2, \"title\": \"Day 2\", \"description\": \"Explore\", \"overnightStop\": \"Hotel\"}, {\"day\": 3, \"title\": \"Day 3\", \"description\": \"Explore\", \"overnightStop\": \"Hotel\"}], \"costs\": {\"total\": \"-10\"}}",
"reason": "non-negative EUR"
},
{
"name": "malformed-cost",
"output": "{\"vehicle\": {\"type\": \"MPV\", \"model\": \"Family MPV\", \"reasoning\": \"Room for luggage\"}, \"routeOverview\": \"Rome\", \"itinerary\": [{\"day\": 1, \"title\": \"Day 1\", \"description\": \"Explore\", \"overnightStop\": \"Hotel\"}, {\"day\": 2, \"title\": \"Day 2\", \"description\": \"Explore\", \"overnightStop\": \"Hotel\"}, {\"day\": 3, \"title\": \"Day 3\", \"description\": \"Explore\", \"overnightStop\": \"Hotel\"}], \"costs\": {\"total\": \"about three hundred\"}}",
"reason": "non-negative EUR"
},
{
"name": "empty-output",
"output": "",
"reason": "no output"
}
]
The file is JSON inside a YAML list so each entry is a single string the strategy can parse. Five defects are pinned: a null vehicle, a gap in the itinerary days, a negative cost, a non-numeric cost, and an empty output.
Create src/test/java/com/tripplanner/evaluation/KnownBadOutputs.java to load the fixtures:
package com.tripplanner.evaluation;
import org.yaml.snakeyaml.LoaderOptions;
import org.yaml.snakeyaml.Yaml;
import java.io.FileInputStream;
import java.io.IOException;
import java.io.InputStream;
import java.util.List;
import java.util.Map;
/**
* Loads the deliberately invalid plan texts from
* {@code src/test/resources/evaluation/known-bad.yaml}. Resolved relative to
* the module working directory, the same way the sample loader does.
*/
public final class KnownBadOutputs {
private static final String PATH = "src/test/resources/evaluation/known-bad.yaml";
private KnownBadOutputs() {
}
@SuppressWarnings("unchecked")
public static List<Map<String, Object>> load() {
Yaml yaml = new Yaml(new LoaderOptions());
try (InputStream in = new FileInputStream(PATH)) {
return yaml.loadAs(in, List.class);
} catch (IOException e) {
throw new IllegalStateException("Unable to load " + PATH, e);
}
}
}
Testing the invariant strategy
Now we can write the test that exercises the strategy against both valid plans and the known-bad fixtures.
Create src/test/java/com/tripplanner/evaluation/TripPlanInvariantStrategyTest.java:
package com.tripplanner.evaluation;
import com.tripplanner.model.TripPlan;
import io.quarkiverse.langchain4j.testing.evaluation.EvaluationSample;
import io.quarkiverse.langchain4j.testing.evaluation.Parameters;
import org.junit.jupiter.api.Test;
import java.util.List;
import java.util.stream.IntStream;
import static org.junit.jupiter.api.Assertions.*;
class TripPlanInvariantStrategyTest {
private final TripPlanInvariantStrategy strategy = new TripPlanInvariantStrategy();
@Test
void acceptsCurrencyAndDecimalTotals() {
for (String total : List.of("310", "€310", "EUR 310.50", "310,50 EUR")) {
var plan = validPlan(3);
var withCost = new TripPlan(plan.vehicle(), plan.routeOverview(), plan.itinerary(),
new TripPlan.CostEstimate("90", "10", "0", "100", "50", "20", total));
assertTrue(strategy.evaluate(sample(3), withCost).passed(), total);
}
}
@Test
void rejectsWrongCountDuplicateAndOutOfRangeDays() {
assertFalse(strategy.evaluate(sample(7), validPlan(1)).passed());
var plan = validPlan(3);
for (List<TripPlan.DayItinerary> days : List.of(
List.of(day(1), day(1), day(3)), List.of(day(0), day(1), day(2)),
List.of(day(1), day(2), day(4)))) {
assertFalse(strategy.evaluate(sample(3),
new TripPlan(plan.vehicle(), plan.routeOverview(), days, plan.costs())).passed());
}
}
@Test
void rejectsMissingFieldsWithoutCrashing() {
var plan = validPlan(3);
for (TripPlan broken : List.of(
new TripPlan(null, plan.routeOverview(), plan.itinerary(), plan.costs()),
new TripPlan(plan.vehicle(), null, plan.itinerary(), plan.costs()),
new TripPlan(plan.vehicle(), plan.routeOverview(), null, plan.costs()),
new TripPlan(plan.vehicle(), plan.routeOverview(), plan.itinerary(), null),
new TripPlan(new TripPlan.VehicleRecommendation("MPV", null, "Room", null),
"Rome", List.of(new TripPlan.DayItinerary(1, "Arrival", null, null)), plan.costs()))) {
assertFalse(strategy.evaluate(sample(3), TripPlanText.render(broken)).passed());
}
assertFalse(strategy.evaluate(sample(3), "null").passed());
assertFalse(strategy.evaluate(sample(3), "not JSON").passed());
}
@Test
void savedPlanRetainsDetailsAndKnownBadFixturesFail() {
assertTrue(TripPlanText.render(validPlan(3)).contains("overnightStop"));
assertTrue(strategy.evaluate(sample(3), TripPlanText.render(validPlan(3))).passed());
for (var bad : KnownBadOutputs.load()) {
var result = strategy.evaluate(sample(3), (String) bad.get("output"));
assertFalse(result.passed(), bad.get("name").toString());
assertTrue(result.explanation().contains((String) bad.get("reason")), result.explanation());
}
}
static TripPlan validPlan(int days) {
return new TripPlan(new TripPlan.VehicleRecommendation("MPV", "Family MPV", "Room for luggage", null),
"Rome", IntStream.rangeClosed(1, days).mapToObj(TripPlanInvariantStrategyTest::day).toList(),
new TripPlan.CostEstimate("€90/day", "€10", "€0", "€100", "€50", "€20", "€310"));
}
static EvaluationSample<String> sample(int days) {
return EvaluationSample.<String>builder().withName("test")
.withParameters(Parameters.of("Rome", "2027-07-10", String.valueOf(days), "family", "2", "moderate", "parks"))
.withExpectedOutput("Follow the request and supplied evidence.").build();
}
private static TripPlan.DayItinerary day(int day) {
return new TripPlan.DayItinerary(day, "Day " + day, "Explore on foot", "Hotel");
}
}
The test builds valid plans with validPlan(), which constructs a complete TripPlan with the right number of days, and checks them against the strategy. It also verifies that currency formats with or without the euro sign pass, that wrong day counts and duplicates fail, that missing fields fail without a crash, and that every entry in known-bad.yaml is caught with the correct reason.
Make sure the Surefire plugin in your pom.xml includes this test:
Run the invariant test:
All four test methods should pass. If an invariant is missing or the fixture format is wrong, the test names the failing entry.
Adding the AI judge
The invariant gate checks structure. It cannot tell you that the plan is for the wrong season, or that the vehicle reasoning contradicts the trip requirements. For that, we need a model to compare the output against the expected output for the same sample.
The judge is a separate model invocation from the planning run. It runs with its own prompt and is never in the planning path, so its usage stays separable from the planner’s measurements.
Create src/test/resources/evaluation/rubric.txt:
You are an AI evaluating a trip plan response and its expected output.
Answer with the single word "true" if the response is an acceptable
equivalent of the expected output for the same trip request, and "false"
otherwise. Do not explain. Do not output anything other than the word
"true" or "false".
Acceptable equivalence means: the same vehicle model, the same number of
itinerary days with the same day numbers, and a cost total that matches to
the unit. Wording, ordering of free-form reasons, and extra explanatory
text may differ.
Response to evaluate: {response}
Expected output: {expected_output}
The rubric asks the model for the single word true or false. AiJudgeStrategy parses the verdict with Boolean.parseBoolean, so a JSON object, a full sentence, or a score with an explanation all read as false.
Create src/test/java/com/tripplanner/evaluation/TripPlanJudge.java:
package com.tripplanner.evaluation;
import dev.langchain4j.model.chat.ChatModel;
import io.quarkiverse.langchain4j.testing.evaluation.EvaluationResult;
import io.quarkiverse.langchain4j.testing.evaluation.EvaluationSample;
import io.quarkiverse.langchain4j.testing.evaluation.judge.AiJudgeStrategy;
import java.nio.file.Files;
import java.nio.file.Path;
import java.util.Map;
/**
* Thin wrapper over {@link AiJudgeStrategy} that loads the rubric from
* {@code src/test/resources/evaluation/rubric.txt} and records which judge
* model was used, so judge usage is separable from application measurements.
*
* <p>The rubric must ask for a literal {@code true}/{@code false} answer:
* {@link AiJudgeStrategy} parses the verdict with
* {@code Boolean.parseBoolean}, so any other shape (JSON, prose) is a failing
* judgment.
*/
public class TripPlanJudge {
private static final Path RUBRIC = Path.of("src/test/resources/evaluation/rubric.txt");
private final ChatModel judgeModel;
private final AiJudgeStrategy strategy;
public TripPlanJudge(ChatModel judgeModel) {
this(judgeModel, loadRubric());
}
public TripPlanJudge(ChatModel judgeModel, String rubric) {
this.judgeModel = judgeModel;
this.strategy = new AiJudgeStrategy(judgeModel, rubric);
}
public EvaluationResult judge(EvaluationSample<String> sample, String actual) {
EvaluationResult result = strategy.evaluate(sample, actual);
// The judge-model metadata is the class of the model used, which is
// stable and does not depend on the model's self-reported name (which
// the ChatModel interface does not expose in this stack).
return result.withMetadata(Map.of("judge-model", judgeModel.getClass().getSimpleName()));
}
private static String loadRubric() {
try {
return Files.readString(RUBRIC);
} catch (Exception e) {
throw new IllegalStateException("Unable to load judge rubric from " + RUBRIC, e);
}
}
}
TripPlanJudge wraps AiJudgeStrategy, loads the rubric from the classpath, and attaches metadata recording which model class was used for the judgment.
Verifying the judge contract offline
The judge contract is worth pinning with a deterministic test. A scripted model always returns a fixed string, so we can verify the parsing rules without any network call: a true verdict passes, a false verdict fails, and a JSON or prose verdict fails.
Create src/test/java/com/tripplanner/evaluation/TripPlanJudgeContractTest.java:
package com.tripplanner.evaluation;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertTrue;
import dev.langchain4j.data.message.AiMessage;
import dev.langchain4j.data.message.UserMessage;
import dev.langchain4j.model.chat.ChatModel;
import dev.langchain4j.model.chat.request.ChatRequest;
import dev.langchain4j.model.chat.response.ChatResponse;
import io.quarkiverse.langchain4j.testing.evaluation.EvaluationResult;
import io.quarkiverse.langchain4j.testing.evaluation.EvaluationSample;
import io.quarkiverse.langchain4j.testing.evaluation.Parameters;
import org.junit.jupiter.api.Test;
import java.util.List;
/**
* The judge contract, verified offline with a deterministic fake model and no
* app boot: a rubric that asks for a literal true/false verdict passes on a
* matching plan and fails otherwise, and a non-boolean or JSON verdict can
* never register as a pass (it parses as false).
*/
class TripPlanJudgeContractTest {
private static final String EXPECTED =
"Vehicle: Family MPV (Room for two travelers and luggage)\n"
+ "Route: Coastal route through the region\n"
+ "Day 1: Arrival\nDay 2: Family day\nDay 3: Departure\n"
+ "Cost total: 310\n";
@Test
void aMatchingPlanPassesTheJudge() {
TripPlanJudge judge = new TripPlanJudge(new JudgeModel(s -> "true"), rubric());
EvaluationResult result = judge.judge(sample(), EXPECTED);
assertTrue(result.passed());
assertEquals(1.0, result.score());
assertEquals("true", result.explanation());
assertEquals("JudgeModel", result.metadata().get("judge-model"));
}
@Test
void aMismatchingPlanFailsTheJudge() {
TripPlanJudge judge = new TripPlanJudge(new JudgeModel(s -> "false"), rubric());
EvaluationResult result = judge.judge(sample(), "A completely different itinerary.");
assertFalse(result.passed());
assertEquals(0.0, result.score());
}
@Test
void aJsonVerdictIsNeverParsedAsAPass() {
// A model that answers JSON (a common mistake) must not pass:
// Boolean.parseBoolean of any non-"true" text is false.
TripPlanJudge judge = new TripPlanJudge(new JudgeModel(s -> "{\"score\": 9, \"explanation\": \"strong match\"}"), rubric());
EvaluationResult result = judge.judge(sample(), EXPECTED);
assertFalse(result.passed());
assertEquals(0.0, result.score());
assertTrue(result.explanation().startsWith("{"));
}
@Test
void proseOrMissingVerdictIsTreatedAsFalse() {
TripPlanJudge judge = new TripPlanJudge(new JudgeModel(s -> "This looks like a good plan overall."), rubric());
EvaluationResult result = judge.judge(sample(), EXPECTED);
assertFalse(result.passed(), "a verdict that is not the word 'true' cannot pass the gate");
}
/**
* A deterministic judge model: ignores the prompt and always returns the
* configured verdict. Proves the parsing contract without any network.
*/
static class JudgeModel implements ChatModel {
private final java.util.function.Function<String, String> verdict;
JudgeModel(java.util.function.Function<String, String> verdict) {
this.verdict = verdict;
}
@Override
public ChatResponse doChat(ChatRequest request) {
String user = request.messages().stream()
.filter(UserMessage.class::isInstance)
.map(UserMessage.class::cast)
.map(UserMessage::singleText)
.reduce("", (a, b) -> a + "\n" + b);
return ChatResponse.builder().aiMessage(AiMessage.from(verdict.apply(user))).build();
}
@Override
public List<dev.langchain4j.model.chat.listener.ChatModelListener> listeners() {
return List.of();
}
}
private static String rubric() {
return "Answer with the single word \"true\" or \"false\".\n\n"
+ "Response to evaluate: {response}\n"
+ "Expected output: {expected_output}\n";
}
private static EvaluationSample<String> sample() {
return EvaluationSample.<String>builder()
.withName("rome-family-three-days")
.withParameters(Parameters.of("Rome", "2027-07-10", "3", "family", "2", "moderate", "coastal towns"))
.withExpectedOutput(EXPECTED)
.build();
}
}
The JudgeModel inner class implements ChatModel and always returns the configured verdict string. The four test methods pin the contract: a matching plan passes, a mismatching plan fails, a JSON verdict is never parsed as a pass, and a prose verdict is treated as false.
Add the test to Surefire:
Run both tests so far:
Recording evaluation runs
A single evaluation run is a snapshot. The useful question is what happens when you run the same sample again with the same inputs and compare the results. We need a way to record each run with its input, saved output, captured evidence, trace id, model name, and per-strategy outcomes.
Create src/test/java/com/tripplanner/evaluation/EvaluationRun.java:
package com.tripplanner.evaluation;
import java.time.Instant;
import java.util.List;
/**
* One evaluated run of the trip planner for a single sample.
*
* @param sampleId the sample name (stable across repeated experiments)
* @param input the request parameters the planner was run with
* @param output the saved plan text (TripPlanText rendering), produced once
* @param evidence supporting evidence captured during the run (e.g. tool results)
* @param traceId the 32-hex OpenTelemetry trace id of the planning run, captured
* while the root span was current; used to attach scores in Langfuse
* @param model the chat model used for the run (class name), for comparability
* @param strategies per-strategy results for the saved output
* @param startedAt when the run started
*/
public record EvaluationRun(
String sampleId,
List<String> input,
String output,
List<String> evidence,
String traceId,
String model,
List<StrategyOutcome> strategies,
Instant startedAt) {
/**
* The outcome of applying one evaluation strategy to the saved output.
*/
public record StrategyOutcome(
String strategy,
boolean passed,
double score,
String reason,
java.util.Map<String, Object> metadata) {
}
/**
* Mean strategy score, clamped to 0..1; 1.0 when no strategy ran.
*/
public double aggregateScore() {
if (strategies.isEmpty()) {
return 1.0;
}
double sum = 0.0;
for (StrategyOutcome s : strategies) {
sum += Math.max(0.0, Math.min(1.0, s.score()));
}
return sum / strategies.size();
}
public boolean allStrategiesPassed() {
return strategies.stream().allMatch(StrategyOutcome::passed);
}
}
EvaluationRun is a record that holds everything about one evaluated run. StrategyOutcome captures the result of one strategy. aggregateScore() returns the mean of the strategy scores, and allStrategiesPassed() is a convenience check.
Create src/test/java/com/tripplanner/evaluation/EvaluationRunRecorder.java:
package com.tripplanner.evaluation;
import java.util.ArrayList;
import java.util.List;
/**
* In-memory store of {@link EvaluationRun}s, keyed by sample id, that keeps
* every repeated experiment so a follow-up run can be compared against the
* first. This is what makes an evaluation repeatable: the same sample run
* again retains comparable inputs, output, and evidence.
*
* <p>Not a Quarkus bean on purpose: it holds the results of a test, not a
* runtime concern, and keeping it plain makes it easy to assert against.
*/
public class EvaluationRunRecorder {
/** All runs for one sample, in submission order. */
public record SampleHistory(String sampleId, List<EvaluationRun> runs) {
public EvaluationRun first() {
return runs.get(0);
}
public EvaluationRun last() {
return runs.get(runs.size() - 1);
}
/** True when a later run's input matches the first run's (comparable experiment). */
public boolean inputsMatch() {
EvaluationRun first = first();
return runs.stream().allMatch(r -> r.input().equals(first.input()));
}
}
private final List<EvaluationRun> runs = new ArrayList<>();
public void record(EvaluationRun run) {
runs.add(run);
}
public List<EvaluationRun> all() {
return List.copyOf(runs);
}
public SampleHistory history(String sampleId) {
List<EvaluationRun> forSample = runs.stream()
.filter(r -> r.sampleId().equals(sampleId))
.toList();
if (forSample.isEmpty()) {
return null;
}
return new SampleHistory(sampleId, List.copyOf(forSample));
}
public int count() {
return runs.size();
}
}
The recorder is an in-memory store that keeps every run per sample, so a follow-up run is compared against the first rather than replacing it. SampleHistory groups the runs for one sample and has an inputsMatch() method that detects when a repeated experiment used different parameters.
Testing the recorder
Create src/test/java/com/tripplanner/evaluation/EvaluationRunRecorderTest.java:
package com.tripplanner.evaluation;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertNull;
import static org.junit.jupiter.api.Assertions.assertSame;
import static org.junit.jupiter.api.Assertions.assertTrue;
import org.junit.jupiter.api.Test;
import java.time.Instant;
import java.util.List;
/**
* The recorder keeps every repeated experiment for a sample so a follow-up run
* can be compared against the first: same inputs, retained output and evidence,
* and per-case strategy results.
*/
class EvaluationRunRecorderTest {
@Test
void aRepeatedExperimentRetainsComparableInputAndEvidence() {
EvaluationRunRecorder recorder = new EvaluationRunRecorder();
String output = "Vehicle: Family MPV\nRoute: Coastal route\nDay 1: Arrival\nCost total: 310\n";
recorder.record(run("rome-family-three-days", output, List.of("weather: sunny", "pois: 3 entries")));
recorder.record(run("rome-family-three-days", output, List.of("weather: sunny", "pois: 3 entries")));
EvaluationRunRecorder.SampleHistory history = recorder.history("rome-family-three-days");
assertNotNull(history);
assertEquals(2, history.runs().size());
assertTrue(history.inputsMatch(), "both runs must have been made from the same request");
// The saved output and evidence are retained across runs for comparison.
assertSame(history.first().output(), history.last().output());
assertEquals(history.first().evidence(), history.last().evidence());
}
@Test
void perCaseResultsAreKeptSeparately() {
EvaluationRunRecorder recorder = new EvaluationRunRecorder();
recorder.record(run("rome", "plan-rome", List.of("e-rome")));
recorder.record(run("barcelona", "plan-barcelona", List.of("e-barcelona")));
assertEquals(2, recorder.count());
assertEquals("plan-rome", recorder.history("rome").first().output());
assertEquals("plan-barcelona", recorder.history("barcelona").first().output());
assertNull(recorder.history("unknown"), "a sample that was never run has no history");
}
@Test
void aMismatchedInputIsDetected() {
EvaluationRunRecorder recorder = new EvaluationRunRecorder();
recorder.record(run("rome", "p", List.of(), List.of("Rome", "3 days")));
recorder.record(run("rome", "p", List.of(), List.of("Rome", "5 days")));
assertFalse(recorder.history("rome").inputsMatch());
}
@Test
void theAggregateScoreIsTheMeanOfStrategies() {
EvaluationRun run = new EvaluationRun(
"rome", List.of("Rome"), "plan", List.of(), "trace-1", "model",
List.of(
new EvaluationRun.StrategyOutcome("invariant", true, 1.0, null, null),
new EvaluationRun.StrategyOutcome("judge", false, 0.0, "mismatch", null)),
Instant.now());
assertEquals(0.5, run.aggregateScore());
assertFalse(run.allStrategiesPassed());
}
private static EvaluationRun run(String sampleId, String output, List<String> evidence) {
return run(sampleId, output, evidence, List.of(sampleId, "3 days"));
}
private static EvaluationRun run(String sampleId, String output, List<String> evidence, List<String> input) {
return new EvaluationRun(
sampleId, input, output, evidence, "trace-" + sampleId, "ScriptedModel", List.of(), Instant.now());
}
}
The test verifies that a repeated experiment retains comparable inputs and evidence, that per-case results stay separate, that mismatched inputs are detected, and that the aggregate score is the mean of the per-strategy scores.
Add the test to Surefire:
Writing evaluation samples
The harness needs a set of known trip requests and their expected outputs. These samples define the inputs that both the offline and live suites use.
Create src/test/resources/evaluation/samples.yaml:
# Evaluation samples for the trip planner (step 07).
#
# The expected-output is the deterministic TripPlanText rendering of the plan
# the scripted model is expected to produce. The live evaluation function
# renders a real run the same way, so a sample passes its invariant and
# judge checks only when the real run matches this shape.
- name: "rome-family-three-days"
parameters:
- "Rome"
- "2027-07-10"
- "3"
- "family"
- "2"
- "moderate"
- "coastal towns"
expected-output: "Vehicle: Family MPV (Room for two travelers and luggage)\nRoute: Coastal route through the region\nDay 1: Arrival\nDay 2: Family day\nDay 3: Departure\nCost total: 310\n"
tags:
- "smoke"
- "mcp"
- name: "barcelona-family-three-days"
parameters:
- "Barcelona"
- "2027-07-12"
- "3"
- "family"
- "2"
- "moderate"
- "beaches"
expected-output: "Vehicle: Family MPV (Room for two travelers and luggage)\nRoute: Coastal route through the region\nDay 1: Arrival\nDay 2: Family day\nDay 3: Departure\nCost total: 310\n"
tags:
- "smoke"
- "mcp"
- name: "rome-family-seven-days"
parameters:
- "Rome"
- "2027-07-10"
- "7"
- "family"
- "2"
- "moderate"
- "historic sites"
expected-output: "Vehicle: Family MPV (Room for two travelers and luggage)\nRoute: Coastal route through the region\nDay 1: Arrival\nDay 2: Family day\nDay 3: Sightseeing\nDay 4: Day trip\nDay 5: Markets\nDay 6: Leisure\nDay 7: Departure\nCost total: 310\n"
tags:
- "smoke"
Each sample has a name, the planner parameters (destination, date, days, travel style, travelers, budget, interests), an expected output string, and tags. The expected output uses the same format as TripPlanText.render() so the invariant and judge checks agree on what they’re comparing.
Assembling the harness test
The harness test ties the samples, the invariant strategy, and the scorer together. It loads samples from the YAML file, runs the invariant strategy against a synthetic valid plan for each sample’s requested duration, and saves a report.
Create src/test/java/com/tripplanner/evaluation/TripPlanEvaluationHarnessTest.java:
package com.tripplanner.evaluation;
import io.quarkiverse.langchain4j.testing.evaluation.EvaluationAssertions;
import io.quarkiverse.langchain4j.testing.evaluation.SampleLoaderResolver;
import io.quarkiverse.langchain4j.testing.evaluation.Scorer;
import org.junit.jupiter.api.Test;
import org.junit.jupiter.api.io.TempDir;
import java.nio.file.Path;
import java.nio.file.Files;
import static org.junit.jupiter.api.Assertions.*;
class TripPlanEvaluationHarnessTest {
@TempDir Path directory;
@Test
void loadsIndependentRequirementsAndChecksEachRequestedDuration() throws Exception {
var samples = SampleLoaderResolver.load("src/test/resources/evaluation/samples.yaml", String.class);
assertEquals(3, samples.size());
var report = new Scorer().evaluate(samples, parameters -> TripPlanText.render(
TripPlanInvariantStrategyTest.validPlan(Integer.parseInt(parameters.get(2).toString()))),
new TripPlanInvariantStrategy());
EvaluationAssertions.assertThat(report).hasAllPassed().hasEvaluationCount(3);
Path file = directory.resolve("report.json");
report.saveAs(file, "json", java.util.Map.of("includeDetails", true));
assertTrue(Files.readString(file).contains("rome-family-seven-days"));
}
@Test
void oneFixedPlanCannotPassRequestsForDifferentDurations() throws Exception {
var samples = SampleLoaderResolver.load("src/test/resources/evaluation/samples.yaml", String.class);
var report = new Scorer().evaluate(samples,
parameters -> TripPlanText.render(TripPlanInvariantStrategyTest.validPlan(1)),
new TripPlanInvariantStrategy());
EvaluationAssertions.assertThat(report).hasEvaluationCount(3).hasFailedCount(3);
}
}
The first test loads the samples, generates a valid plan for each one, runs the invariant strategy, and asserts all pass. It also saves the report as JSON and checks that one of the sample names (rome-family-seven-days) appears in it. The second test proves that a fixed one-day plan cannot pass samples requesting different durations, confirming the harness actually catches the mismatch.
Add the test to Surefire:
Running the offline suite
The offline suite runs all four test classes against the fixtures, with no model provider, no container runtime, and no network.
Run the default test suite in section-3/step-07/trip-planner:
Your Surefire <includes> should now list all four tests:
<includes>
<include>**/TripPlanInvariantStrategyTest.java</include>
<include>**/TripPlanJudgeContractTest.java</include>
<include>**/TripPlanEvaluationHarnessTest.java</include>
<include>**/EvaluationRunRecorderTest.java</include>
</includes>
The suite loads the samples from samples.yaml, applies the invariant strategy to known-good and known-bad outputs, and verifies the judge contract and the recorder. When you change a prompt, a skill, or a model, run this suite first. It is fast enough to run after every edit, and a red invariant check on a known-good sample means the rendering or the plan shape changed in a way the gate now detects.
Configuring telemetry for the live runs
Because quarkus-langfuse and quarkus-opentelemetry are runtime dependencies, Langfuse Dev Services starts automatically whenever you run the application — in dev mode, traces appear in Langfuse as planning runs complete. You can find the UI URL, login credentials, and all injected configuration in the Dev UI under Dev Services. You may notice that startup takes longer than in previous steps: Dev Services is now pulling and starting containers for Kafka, PostgreSQL, and Langfuse. In production none of those would be managed by Quarkus, so startup time would be unaffected.
The Extensions page also has a Quarkus Langfuse card with a direct link to the Langfuse UI, and an Observability card for the OpenTelemetry stack.
The tests need a bit more care. The offline suite runs with no containers and no network, so the %test profile disables Dev Services and points the OTLP exporter at a dead endpoint. The %mcp profile does the same because the composition IT checks plan structure rather than telemetry. The %evals profile, by contrast, keeps Dev Services on (it is already the default), marks the container as shared so Failsafe reuses the one already running rather than starting a fresh instance, and widens the span filter so the planning root span — which carries no gen_ai attributes — is not dropped from the trace.
Add the following to your src/test/resources/application.properties:
# Offline test profile: Langfuse Dev Services and OTLP are disabled so the
# fast offline suite needs no containers and no network.
%test.quarkus.langfuse.devservices.enabled=false
%test.quarkus.otel.exporter.otlp.protocol=http/protobuf
%test.quarkus.otel.exporter.otlp.endpoint=http://127.0.0.1:9/v1
%test.quarkus.langfuse.base-url=http://127.0.0.1:9
%test.quarkus.langfuse.public-key=pk-offline
%test.quarkus.langfuse.secret-key=sk-offline
# Live composition (-Pevals, @TestProfile mcp): the MCP workflow runs against
# the real Trip Intelligence server; Langfuse stays off because this IT
# asserts plan structure, not telemetry.
%mcp.quarkus.langfuse.devservices.enabled=false
%mcp.quarkus.langfuse.base-url=http://127.0.0.1:9
%mcp.quarkus.langfuse.public-key=pk-offline
%mcp.quarkus.langfuse.secret-key=sk-offline
%mcp.quarkus.otel.exporter.otlp.enabled=false
# Live evaluation (-Pevals, @TestProfile evals): Langfuse Dev Services is
# already on by default; shared=true reuses the running container rather than
# starting a new one. span-filter=ALL exports the planning root span, which
# has no gen_ai attributes and would otherwise be dropped.
%evals.quarkus.langfuse.devservices.shared=true
%evals.quarkus.langfuse.otel.span-filter=ALL
Adding the evals Maven profile
The live evaluation tests run under a separate Maven profile so ./mvnw test never triggers them. The profile selects only the *LiveIT classes through Failsafe.
Add the evals profile to your trip-planner/pom.xml, inside the <profiles> block:
<profile>
<!--
Live evaluation and composition. Runs only the *LiveIT classes
through Failsafe. Requires a running Trip Intelligence MCP server
on :8085, a container runtime for Langfuse Dev Services (quality
IT), and a real LLM key. The default ./mvnw test never runs these.
-->
<id>evals</id>
<properties>
<skipITs>false</skipITs>
</properties>
<build>
<plugins>
<plugin>
<artifactId>maven-failsafe-plugin</artifactId>
<version>${surefire-plugin.version}</version>
<configuration>
<argLine>@{argLine}</argLine>
<forkedProcessTimeoutInSeconds>1800</forkedProcessTimeoutInSeconds>
<includes>
<include>**/*LiveIT.java</include>
</includes>
<systemPropertyVariables>
<java.util.logging.manager>org.jboss.logmanager.LogManager</java.util.logging.manager>
<maven.home>${maven.home}</maven.home>
</systemPropertyVariables>
</configuration>
</plugin>
</plugins>
</build>
</profile>
Live evaluation: the composition run
The composition run drives the full planning graph with a scripted model while the two @McpClientAgent subagents call the real Trip Intelligence server. No LLM is called, so it is cheap to run, but it exercises the argument flow, the MCP data path, and the request isolation of the complete workflow. Each saved output is checked by the invariant strategy and recorded.
Create src/test/java/com/tripplanner/evaluation/TripPlannerCompositionLiveIT.java:
package com.tripplanner.evaluation;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
import com.tripplanner.agentic.workflow.TripPlannerSystem;
import com.tripplanner.model.ItineraryResult;
import com.tripplanner.model.TripPlan;
import com.tripplanner.model.TripPlan.DayItinerary;
import com.tripplanner.model.TripPlan.VehicleRecommendation;
import com.tripplanner.model.VehicleEvaluation;
import dev.langchain4j.data.message.AiMessage;
import dev.langchain4j.data.message.UserMessage;
import dev.langchain4j.model.chat.ChatModel;
import dev.langchain4j.model.chat.request.ChatRequest;
import dev.langchain4j.model.chat.response.ChatResponse;
import io.quarkiverse.langchain4j.ModelName;
import io.quarkiverse.langchain4j.testing.evaluation.EvaluationSample;
import io.quarkiverse.langchain4j.testing.evaluation.Parameters;
import io.quarkus.test.junit.QuarkusTest;
import io.quarkus.test.junit.QuarkusTestProfile;
import io.quarkus.test.junit.TestProfile;
import jakarta.enterprise.inject.Alternative;
import jakarta.inject.Inject;
import jakarta.inject.Singleton;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.time.Instant;
import java.util.List;
import java.util.Set;
import java.util.concurrent.atomic.AtomicReference;
/**
* Drives the final planning graph with a scripted model while the two
* {@code @McpClientAgent} subagents call the real Trip Intelligence MCP server
* over streamable HTTP at http://localhost:8085/mcp. One fresh run per sample;
* each saved output is checked by the invariant strategy and retained in the
* recorder, so a repeated experiment keeps comparable inputs and evidence.
*
* <p>Run with {@code ./mvnw verify -Pevals} after starting the MCP server
* (sunny fixture). No LLM is called (scripted model) and no Langfuse
* container is started (dead keys under the {@code mcp} profile).
*/
@QuarkusTest
@TestProfile(TripPlannerCompositionLiveIT.ScriptedProfile.class)
class TripPlannerCompositionLiveIT {
@Inject
TripPlannerSystem tripPlannerSystem;
@Inject
ChatModel chatModel;
ScriptedProfile.ScriptedModel model;
final EvaluationRunRecorder recorder = new EvaluationRunRecorder();
final TripPlanInvariantStrategy invariants = new TripPlanInvariantStrategy();
@BeforeEach
void reset() {
model = (ScriptedProfile.ScriptedModel) chatModel;
// Skip where no Trip Intelligence server is running the sunny fixture,
// so the suite stays green in environments without the MCP server.
org.junit.jupiter.api.Assumptions.assumeTrue(
probeSunnyServer(),
"Trip Intelligence MCP server on :8085 is not running the sunny fixture");
}
static boolean probeSunnyServer() {
try {
var client = java.net.http.HttpClient.newHttpClient();
var init = client.send(
mcpRequest("initialize",
"{\"protocolVersion\":\"2025-11-25\",\"capabilities\":{},\"clientInfo\":{\"name\":\"it\",\"version\":\"1\"}}")
.build(),
java.net.http.HttpResponse.BodyHandlers.ofString());
String session = init.headers().firstValue("mcp-session-id").orElse(null);
var call = mcpRequest("tools/call",
"{\"name\":\"getWeatherForecast\",\"arguments\":{\"destination\":\"Rome\",\"startDate\":\"2027-07-10\",\"days\":\"3\"}}");
if (session != null) {
call.header("mcp-session-id", session);
}
String body = client.send(call.build(), java.net.http.HttpResponse.BodyHandlers.ofString()).body();
return body != null && body.contains("Mostly sunny");
} catch (Exception e) {
return false;
}
}
static java.net.http.HttpRequest.Builder mcpRequest(String method, String params) {
return java.net.http.HttpRequest.newBuilder(java.net.URI.create("http://localhost:8085/mcp"))
.timeout(java.time.Duration.ofSeconds(3))
.header("Content-Type", "application/json")
.header("Accept", "application/json, text/event-stream")
.POST(java.net.http.HttpRequest.BodyPublishers.ofString(
"{\"jsonrpc\":\"2.0\",\"id\":1,\"method\":\"" + method + "\",\"params\":" + params + "}"));
}
@Test
void aRomeFamilyTripPassesItsInvariantsAndIsRecorded() {
TripPlan plan = tripPlannerSystem.planTrip("Rome", "2027-07-10", "3", "family", "2", "moderate", "coastal towns");
String output = TripPlanText.render(plan);
// The real MCP fixture data reached the research agents.
assertTrue(model.itineraryPrompt().contains("Workshop fixture: 3 days in Rome"),
"real weather fixture must reach the itinerary planner");
assertTrue(model.itineraryPrompt().contains("Rome Zoo & Aquarium"),
"seeded POIs must reach the itinerary planner");
// The saved output passes the deterministic gate.
var invariant = invariants.evaluate(sample("rome-family-three-days", output), output);
assertTrue(invariant.passed(), "saved plan must pass the invariants: " + invariant.explanation());
// The run is retained for a comparable repeated experiment.
var run = new EvaluationRun(
"rome-family-three-days",
List.of("Rome", "2027-07-10", "3", "family", "2", "moderate", "coastal towns"),
output,
List.of(model.evidenceLine()),
null, "ScriptedModel",
List.of(new EvaluationRun.StrategyOutcome("invariant",
invariant.passed(), invariant.score(), invariant.explanation(), invariant.metadata())),
Instant.now());
recorder.record(run);
assertEquals(1, recorder.count());
assertTrue(recorder.history("rome-family-three-days").inputsMatch());
}
@Test
void requestIsolationMeansASecondDestinationIsIndependent() {
TripPlan first = tripPlannerSystem.planTrip("Rome", "2027-07-10", "3", "family", "2", "moderate", "coastal towns");
TripPlan second = tripPlannerSystem.planTrip("Barcelona", "2027-07-12", "3", "family", "2", "moderate", "beaches");
// Each request carried its own destination through the MCP tools.
assertTrue(model.itineraryPrompt().contains("Barcelona"),
"the second request's destination must reach the planner");
assertEquals(3, first.itinerary().size());
assertEquals(3, second.itinerary().size());
assertNotNull(first.vehicle());
assertNotNull(second.vehicle());
assertFalse(TripPlanText.render(first).contains("Barcelona"),
"the first plan must not leak the second request's destination");
}
static EvaluationSample<String> sample(String name, String output) {
return EvaluationSample.<String>builder()
.withName(name)
.withParameters(Parameters.of("Rome", "2027-07-10", "3", "family", "2", "moderate", "coastal towns"))
.withExpectedOutput(output)
.build();
}
public static class ScriptedProfile implements QuarkusTestProfile {
@Override
public String getConfigProfile() {
return "mcp";
}
@Override
public Set<Class<?>> getEnabledAlternatives() {
return Set.of(ScriptedModel.class, EnhancedModel.class);
}
@Alternative
@Singleton
@ModelName("enhancedModel")
public static class EnhancedModel implements ChatModel {
@Inject
ChatModel delegate;
@Override
public ChatResponse doChat(ChatRequest request) {
return delegate.doChat(request);
}
}
@Alternative
@Singleton
public static class ScriptedModel implements ChatModel {
@Inject
com.fasterxml.jackson.databind.ObjectMapper mapper;
final AtomicReference<String> itineraryPrompt = new AtomicReference<>();
final AtomicReference<String> lastEvidence = new AtomicReference<>();
String itineraryPrompt() {
return itineraryPrompt.get();
}
String evidenceLine() {
return lastEvidence.get();
}
@Override
public ChatResponse doChat(ChatRequest request) {
String prompt = request.messages().reversed().stream().filter(UserMessage.class::isInstance)
.map(UserMessage.class::cast).map(UserMessage::singleText)
.filter(s -> s.contains("- Destination:") || s.contains("Vehicle:")
|| s.contains("Current recommendation:"))
.findFirst().orElseThrow();
Object result;
if (prompt.contains("vehicle specialist")) {
result = new VehicleRecommendation("MPV", "Family MPV", "Room for two travelers and luggage", null);
} else if (prompt.contains("itinerary planner")) {
itineraryPrompt.set(prompt);
result = new ItineraryResult("Coastal route through the region", List.of(
new DayItinerary(1, "Arrival", "Settle in and explore the old town on foot", "Hotel"),
new DayItinerary(2, "Family day", "Visit the zoo and the science museum", "Hotel"),
new DayItinerary(3, "Departure", "Morning in the park, then the drive home", "Hotel")));
} else if (prompt.contains("evaluator for")) {
lastEvidence.set("evaluation: 8.5");
result = new VehicleEvaluation(8.5, "Keep the MPV");
} else if (prompt.contains("recommendation specialist")) {
throw new AssertionError("The sunny initial-pass script must not revise the vehicle");
} else if (prompt.contains("cost estimation")) {
result = new TripPlan.CostEstimate("90", "10", "0", "100", "50", "20", "310");
} else {
throw new AssertionError("Unexpected agent prompt");
}
return ChatResponse.builder().aiMessage(AiMessage.from(mapper.valueToTree(result).toString())).build();
}
}
}
}
The test uses a ScriptedProfile that activates the mcp config profile and substitutes a ScriptedModel that returns deterministic responses for each agent. The @BeforeEach method probes the MCP server on port 8085 and skips the test if it isn’t running the sunny fixture, so the suite stays green in environments without the server.
aRomeFamilyTripPassesItsInvariantsAndIsRecorded() plans a Rome trip, verifies that the real weather and POI data from the MCP server reached the itinerary planner’s prompt, runs the invariant strategy on the saved output, and records the result. requestIsolationMeansASecondDestinationIsIndependent() plans two trips in sequence and checks that the first plan doesn’t leak the second request’s destination.
Start the MCP server in section-3/step-07/mcp-server, then run the composition IT:
./mvnw -f section-3/step-07/pom.xml -pl trip-planner -Pevals -Dit.test=TripPlannerCompositionLiveIT -Dquarkus.http.test-port=0 verify
The test asserts that the real fixture data reached the research agents, that the saved plan passes the invariants, and that a second request for a different destination is independent of the first.
Publishing scores to Langfuse
Before we write the quality evaluation test, we need a way to publish evaluation scores to Langfuse and verify they land on the correct trace.
Create src/test/java/com/tripplanner/evaluation/LangfuseScorePublisher.java:
package com.tripplanner.evaluation;
import io.quarkiverse.langfuse.api.LangfuseOperations;
import io.quarkiverse.langfuse.api.ScoreFilter;
import com.langfuse.api.model.CreateScoreRequest;
import com.langfuse.api.model.CreateScoreResponse;
import com.langfuse.api.model.CreateScoreSource;
import com.langfuse.api.model.CreateScoreValue;
import com.langfuse.api.model.ScoreDataType;
import java.util.Map;
/**
* Publishes a run's aggregate score to Langfuse, attached to the trace the run
* actually produced (captured while the root span was current).
*
* <p>Kept as a plain class that receives the injected {@link LangfuseOperations}:
* the harness must hold the injected reference, not look the bean up at runtime
* (Arc bean removal drops a bean that is only referenced through
* {@code CDI.current().select()} in a test lambda).
*/
public class LangfuseScorePublisher {
public record PublishedScore(String scoreId, String traceId, String name, double value) {
}
private final LangfuseOperations langfuse;
public LangfuseScorePublisher(LangfuseOperations langfuse) {
this.langfuse = langfuse;
}
/**
* Attaches a numeric score to the given trace and returns the created score
* id for later verification through the score API.
*/
public PublishedScore publishScore(EvaluationRun run, double score, String name, String comment) {
CreateScoreResponse created = langfuse.scores().create(CreateScoreRequest.builder()
.name(name)
.value(new CreateScoreValue(score))
.dataType(ScoreDataType.NUMERIC)
.traceId(run.traceId())
.source(CreateScoreSource.API)
.comment(comment)
.metadata(Map.of(
"sample.id", run.sampleId(),
"model", run.model() == null ? "unknown" : run.model()))
.build());
return new PublishedScore(created.getId(), run.traceId(), name, score);
}
/**
* Returns true when a previously created score is queryable for the run's
* trace. Score ingestion is asynchronous, so callers must retry (e.g. with
* Awaitility).
*/
public boolean scoreIsQueryable(String scoreId, String traceId) {
return langfuse.scores()
.matching(ScoreFilter.builder().traceId(traceId).build())
.findAll().stream()
.map(s -> s.getNumericScoreV31() == null ? null : s.getNumericScoreV31().getId())
.anyMatch(scoreId::equals);
}
}
publishScore() creates a numeric score on the run’s trace id with metadata carrying the sample id and model name. scoreIsQueryable() checks whether the score has been ingested and is queryable on a given trace. Score ingestion in Langfuse is asynchronous, so callers must retry rather than assuming the score is queryable immediately after creation.
Live evaluation: the quality run
The quality run performs one real planning run with your configured model, captures the planning trace while the root span is current, applies the invariant strategy and the judge to the saved output, and publishes the resulting score to Langfuse. It then verifies through the score API that the score landed on exactly the trace of the run it measured.
Create src/test/java/com/tripplanner/evaluation/TripPlanQualityEvaluationLiveIT.java:
package com.tripplanner.evaluation;
import static org.junit.jupiter.api.Assertions.assertEquals;
import static org.junit.jupiter.api.Assertions.assertFalse;
import static org.junit.jupiter.api.Assertions.assertNotNull;
import static org.junit.jupiter.api.Assertions.assertTrue;
import com.tripplanner.agentic.workflow.TripPlannerSystem;
import com.tripplanner.model.TripPlan;
import dev.langchain4j.model.chat.ChatModel;
import io.opentelemetry.api.trace.Span;
import io.opentelemetry.api.trace.SpanKind;
import io.opentelemetry.api.trace.Tracer;
import io.quarkiverse.langfuse.api.LangfuseOperations;
import io.quarkiverse.langchain4j.testing.evaluation.EvaluationResult;
import io.quarkiverse.langchain4j.testing.evaluation.EvaluationSample;
import io.quarkiverse.langchain4j.testing.evaluation.Parameters;
import io.quarkus.test.junit.QuarkusTest;
import io.quarkus.test.junit.QuarkusTestProfile;
import io.quarkus.test.junit.TestProfile;
import jakarta.inject.Inject;
import org.awaitility.Awaitility;
import org.junit.jupiter.api.BeforeEach;
import org.junit.jupiter.api.Test;
import java.time.Duration;
import java.time.Instant;
import java.util.List;
import java.util.Set;
/**
* The full live evaluation loop: one real planning run (real LLM plus the real
* Trip Intelligence MCP server on :8085), its trace captured while the root
* span is current, the saved output checked by the deterministic invariant
* strategy and by the AI judge, and the resulting score published to Langfuse
* and verified back through the score API as landing on exactly this trace.
*
* <p>Run with {@code ./mvnw verify -Pevals}. Requires a container runtime
* (Langfuse Dev Services) and a real LLM key. The MCP server must be running.
*/
@QuarkusTest
@TestProfile(TripPlanQualityEvaluationLiveIT.LiveProfile.class)
@org.junit.jupiter.api.condition.EnabledIf("containerRuntimeAvailable")
class TripPlanQualityEvaluationLiveIT {
static boolean containerRuntimeAvailable() {
try {
Process p = new ProcessBuilder("bash", "-c",
"command -v podman >/dev/null 2>&1 || command -v docker >/dev/null 2>&1")
.redirectErrorStream(true).start();
return p.waitFor() == 0;
} catch (Exception e) {
return false;
}
}
private static final String SAMPLE = "rome-family-three-days";
@Inject
TripPlannerSystem tripPlannerSystem;
@Inject
ChatModel chatModel;
@Inject
Tracer tracer;
@Inject
LangfuseOperations langfuse;
final EvaluationRunRecorder recorder = new EvaluationRunRecorder();
final TripPlanInvariantStrategy invariants = new TripPlanInvariantStrategy();
@Test
void oneSampleOneTraceOneScoreVerifiedThroughTheScoreApi() {
org.junit.jupiter.api.Assumptions.assumeTrue(
TripPlannerCompositionLiveIT.probeSunnyServer(),
"Trip Intelligence MCP server on :8085 is not running the sunny fixture");
// One real planning run, with the root span current so the agent's
// internal spans join this trace. Capture the trace id before ending.
String traceId;
TripPlan plan;
Span root = tracer.spanBuilder("trip-plan-evaluation")
.setSpanKind(SpanKind.INTERNAL)
.setAttribute("trip.sample.id", SAMPLE)
.startSpan();
try (var scope = root.makeCurrent()) {
traceId = root.getSpanContext().getTraceId();
plan = tripPlannerSystem.planTrip("Rome", "2027-07-10", "3", "family", "2", "moderate", "coastal towns");
} finally {
root.end();
}
assertTrue(traceId.matches("[0-9a-f]{32}"), "captured trace id must be 32-hex: " + traceId);
String output = TripPlanText.render(plan);
// Deterministic gate on the saved output.
EvaluationResult invariant = invariants.evaluate(sample(output), output);
// The AI judge is a separate invocation from the application run, so its
// usage stays separable from the planner's measurements.
TripPlanJudge judge = new TripPlanJudge(chatModel);
EvaluationResult judged = judge.judge(sample(output), output);
assertNotNull(judged.metadata().get("judge-model"), "judge must record which model it used");
var run = new EvaluationRun(
SAMPLE,
List.of("Rome", "2027-07-10", "3", "family", "2", "moderate", "coastal towns"),
output,
List.of("vehicle=" + plan.vehicle().model(), "days=" + plan.itinerary().size()),
traceId, "openai",
List.of(
new EvaluationRun.StrategyOutcome("invariant", invariant.passed(), invariant.score(), invariant.explanation(), invariant.metadata()),
new EvaluationRun.StrategyOutcome("judge", judged.passed(), judged.score(), judged.explanation(), judged.metadata())),
Instant.now());
recorder.record(run);
// Publish the score to the captured trace and verify it lands there.
LangfuseScorePublisher publisher = new LangfuseScorePublisher(langfuse);
LangfuseScorePublisher.PublishedScore published =
publisher.publishScore(run, run.aggregateScore(), "plan-quality",
"invariant=" + invariant.passed() + " judge=" + judged.passed());
// Score ingestion is asynchronous, so retry.
Awaitility.await().atMost(Duration.ofSeconds(90)).untilAsserted(() ->
assertTrue(publisher.scoreIsQueryable(published.scoreId(), published.traceId()),
"the score must be queryable on the planning trace"));
// A score must land only on the intended trace, not a foreign one.
String foreignTrace = "0".repeat(32);
assertFalse(publisher.scoreIsQueryable(published.scoreId(), foreignTrace),
"the score must not be queryable on an unrelated trace");
assertEquals(published.traceId(), run.traceId());
}
static EvaluationSample<String> sample(String output) {
return EvaluationSample.<String>builder()
.withName(SAMPLE)
.withParameters(Parameters.of("Rome", "2027-07-10", "3", "family", "2", "moderate", "coastal towns"))
.withExpectedOutput(output)
.build();
}
public static class LiveProfile implements QuarkusTestProfile {
@Override
public String getConfigProfile() {
return "evals";
}
@Override
public Set<Class<?>> getEnabledAlternatives() {
return Set.of();
}
}
}
The test opens a root span around the planning call so the agent invocations create child spans on the same trace. It captures the trace id while the span is current, because the id is only stable while the span is open. After the planning run, it renders the saved output, runs the invariant gate, runs the AI judge, builds an EvaluationRun, publishes the aggregate score to Langfuse, and uses Awaitility to retry until the score is queryable on the correct trace. A final check confirms the score is not queryable on an unrelated trace.
The test is gated by @EnabledIf("containerRuntimeAvailable") so it skips cleanly when Docker or Podman is not installed.
Run the quality IT with a container runtime available:
./mvnw -f section-3/step-07/pom.xml -pl trip-planner -Pevals -Dit.test=TripPlanQualityEvaluationLiveIT -Dquarkus.http.test-port=0 verify
The first run pulls the Langfuse Dev Services images, which is a several-hundred-megabyte download. Subsequent runs reuse them.
What to look for in Langfuse
With the quality run complete, the Langfuse UI shows the planning trace with its spans and the attached plan-quality score. The score’s metadata carries the sample id and the model name, so you can filter by sample or by model when you run the suite several times.
Scroll down to see trace latencies, generation times, and observation spans. The tooltip on a trace name shows the full agent class and method, so you can match a slow span to a specific agent in your code.
Use the trace to read what the run actually did. The LLM spans show the prompts and responses, and the MCP spans show the tool calls and their results. If a sample fails its judge check, the trace is where you find out whether the model ignored the MCP data, reasoned from it, or produced a plan the judge could not match to the expected output.
What’s next?
The trip planner is now backed by a repeatable evaluation harness: deterministic invariant checks catch structural problems without calling a model, a judge measures qualitative fitness against a rubric, and every live run produces a trace in Langfuse with its evaluation score attached. That’s the end of Section 3. Head to the conclusion for a recap of the enterprise patterns you’ve built across all seven steps.



