Skip to content
llm-speed

Gemma 4 E2B JSON output: valid JSON can still route incorrectly

Published 2026-09-15

A JSON schema fixed formatting in our small Gemma 4 E2B test, but it did not guarantee the right decision. On September 15, 2026, we sent nine short helper tasks to Google's Gemma 4 E2B QAT Q4_0 in llama.cpp on an RTX 5090. Each task ran once with a prompt-only JSON instruction and once with a schema constraint.

With the schema, all nine replies parsed as JSON and eight matched the expected values. The remaining reply was valid JSON with an allowed, but incorrect, routing value. This is an exploratory integration test with complete evidence, not a representative model-quality benchmark or a signed leaderboard run.

Download all 18 requests, responses, expected answers and settings (JSON).

What changed when we added a JSON schema?

Nine matched helper inputs, one reply per condition
ConditionStrict JSON parsesParses and matches expected values
Prompt says to return JSON0 of 90 of 9
Same prompt plus schema constraint9 of 98 of 9

The prompt-only replies included Markdown code fences, which fail direct json.loads parsing. Some contained correct values inside those fences; the zero is a strict integration result, not a claim that every answer was semantically wrong. We did not strip fences or repair replies before scoring.

The runtime supports schema-constrained output through response_format. See the official llama.cpp server documentation. The routing schema used an object with one required route field, limited to code, lookup or clarify; additional properties were disallowed.

"response_format": {
  "type": "json_object",
  "schema": {
    "type": "object",
    "properties": {
      "route": { "type": "string", "enum": ["code", "lookup", "clarify"] }
    },
    "required": ["route"],
    "additionalProperties": false
  }
}

The failure a schema could not prevent

The routing instruction assigned programming requests to code, factual questions to lookup, and absent requests to clarify. For Explain the difference between RAM and VRAM., the expected route was lookup. The constrained reply was:

{ "route": "code" }

The reply obeys the schema. It still sends the request to the wrong handler under our stated routing policy. Before giving a helper responsibility for tools or workflow decisions, check both its output structure and whether its decisions match your application's rules. Keep an explicit fallback for cases the helper cannot handle reliably.

Exact model, setup and scope

We tested Google's pinned E2B QAT Q4_0 GGUF, file gemma-4-E2B_q4_0-it.gguf. Its full SHA-256 and byte size are in the download. The runtime was llama.cpp commit 9725a313be0528214c4a02fed906ddaf7b3f712e, running under WSL Linux on an RTX 5090.

Settings were 4,096-token context, one slot, all-layer offload requested, flash attention on, batch size 128, microbatch 64, thinking off, temperature 0, seed 42, output cap 96, and prompt/RAM prompt caching off. The download records the complete request bodies and execution order, shuffled with seed 42. These short inputs do not test the model's full context capacity.

The coding and planner services remained loaded and idle during checks. The runtime reported all 36 layers offloaded, while retaining a CPU-mapped model buffer of about 2,153 MiB. An all-layer offload message therefore does not establish that every model allocation resides in VRAM. This test is not an isolated speed or peak-memory comparison.

The nine inputs cover routing, extracting a model name and memory amount, and leaving an absent price as null. They are hand-written, include two blank-request variants, and are not a representative agent evaluation. Each condition ran once per input; three earlier smoke replies are excluded from the table. We tested neither tools, multimodal input, other quantizations, nor thinking-enabled settings. The results do not establish that E2B is generally suitable or unsuitable for agents.

Why check this before benchmarking a helper?

We started with a practical question: can a small helper share a GPU with a main coding model? E2B loaded, but the routing failure means latency alone would not settle that choice. We deferred the concurrency comparison until the intended helper tasks have a useful success check.

Use our multiple-model residency and queue checklist for the hardware side. For a separate, signed speed study of larger current models, see Qwen3.8-27B versus Gemma 4 12B on the RTX 5090. Those speed measurements do not answer this helper-quality question.