Skip to Content

GPT-6 Astra & Claude Opus 5.5: Valid JSON, Wrong Answer

September 30, 2026 by
GPT-6 Astra & Claude Opus 5.5: Valid JSON, Wrong Answer
Deny Smith

A successful parse tells you that the response can be read as JSON. It doesn't tell you whether the quantity belongs to the item you asked about.

JSON extraction errors can survive schema validation because a well-formed response may still describe the wrong entity. When comparing GPT-6 Astra and Claude Opus 5.5 for document extraction, inspect the relationship between a value and its source, rather than treating a successful parse as a completed job.

OpenAI's structured-output documentation explicitly notes that structured results can contain mistakes. A schema can constrain the shape of an answer; the application still needs a way to check whether the answer corresponds to the document.

1. Reproduce the JSON extraction error

Start a trial with GPT-6 Astra using a small synthetic fixture. Two similar rows can reveal a mistake more clearly than a large production document. The following example is invented and contains no customer data:

Item ID

Description

Quantity

Unit

P-17

Desk lamp, green

6

pieces

P-71

Desk lamp, blue

9

pieces

The task is to extract the quantity for P-17. This hypothetical response has a plausible structure:

{

"item_id": "P-17",

"quantity": 9,

"unit": "pieces"

}

It is also wrong. The quantity came from P-71. A schema requiring a string identifier, an integer quantity and a permitted unit could accept all three fields without detecting the swap.

Keep the input and the erroneous response together. The model settings are worth recording, but first confirm that the application actually sent the row labels and column headings. If a PDF conversion dropped the identifier column, the error begins before the model call.

2. Bind every value to an entity

Ask for source evidence alongside extracted values. That evidence might include a page identifier, a row identifier and the exact text used for the field. Choose a form the application can check against the input it supplied.

For the synthetic table, an evidence-bearing candidate might look like this:

{

"item_id": "P-17",

"quantity": 6,

"unit": "pieces",

"source_row_id": "row-1",

"source_text": "P-17 | Desk lamp, green | 6 | pieces"

}

The example is an application design, not a provider-specific request format. Its value comes from giving the reviewer a short path back to the relevant row.

The source text itself still needs checking. A model can return an invented quotation or a row identifier that doesn't exist. Verify that the cited span occurs in the supplied input, then check that it supports the extracted value and belongs to the requested entity.

Exact occurrence alone isn't enough. The blue lamp's row also occurs in the input. The entity match is what prevents that valid span from supporting the green lamp's quantity.

3. Validate meaning outside the model

Separate validation into checks the program can perform. Parse the response and validate its schema first. Then look up the item identifier in the source and confirm that the cited row belongs to that item. Finally, compare the quantity with the value under the quantity heading. Passing the first check says nothing about the last one.

For a clean table, some of these checks can be deterministic. If the input has already been parsed into reliable rows, a direct lookup may be preferable to asking a language model to recover the value at all. Reserve model work for the ambiguity that requires it.

Documents with merged cells or damaged OCR need more care. A neighbouring label may apply across several rows, and a missing character can change an identifier. Preserve enough surrounding material for a reviewer to inspect that relationship. A narrowly clipped number without its heading loses the evidence needed to resolve the question.

Use an explicit review state when a check fails or the source is unclear. Don't convert missing values to zero merely because the destination field is numeric. “Unknown” and “none” can describe different situations, and the application should preserve that distinction.

Keep business rules separate from extraction rules. A quantity outside an expected range may deserve review, but that doesn't prove the document was read incorrectly. The source itself may contain an unusual but valid value.

Normalisation deserves a separate record as well. Suppose a source says “2 boxes of 6” and the destination expects individual pieces. Returning 12 may be the desired transformation, but 12 is not a verbatim value in the source. Store the raw expression and identify the conversion so a reviewer can distinguish arithmetic from extraction.

Similarly, preserve identifiers as identifiers. If a product code is “0017,” converting it to the number 17 may discard meaningful information. Decide the field's type from its role in the document, and make that decision before building the schema. More elaborate prompting won't repair a destination field that cannot represent the value correctly.

4. Give the second model the evidence

A comparison run with Claude Opus 5.5 should receive the original fixture and the extraction specification. Let it produce its own candidate before showing it the first answer. Then compare field-level results and source references.

Where the models disagree, resolve the field against the document. Where they agree, apply the same validation. Two identical values can come from the same misleading row layout.

For a targeted review, you can give a model the disputed field and its surrounding source, then ask it to explain which label governs the value. Keep the explanation outside the machine-ingested record until the application or a reviewer accepts the correction.

Measure the error that matters to the task. A record with twenty correct fields and one wrong account identifier may be unusable. An average field-accuracy figure alone can obscure that failure, so identify fields whose errors should reject the entire record.

5. Keep the correction in the test set

Once the row swap is resolved, retain a small regression fixture with the expected answer. Include a similar case whose identifiers differ by one character, plus a case with a genuinely missing value. These examples exercise distinct failure conditions.

When a prompt or model changes, rerun the fixtures and inspect any newly accepted errors. A revision that produces prettier JSON but swaps the same source rows hasn't fixed the underlying problem.

JSON extraction errors become manageable when a record can explain where its values came from. Parsing is the first checkpoint. The record is ready to use when the values, entities and source evidence agree.



 

GPT-6 Astra & Claude Opus 5.5: Valid JSON, Wrong Answer
Deny Smith September 30, 2026

Lewis Calvert is the Founder and Editor of Big Write Hook, focusing on digital journalism, culture, and online media. He has 6 years of experience in content writing and marketing and has written and edited many articles on news, lifestyle, travel, business, and technology. Lewis studied Journalism and works to publish clear, reliable, and helpful content while supporting new writers on the Big Write Hook platform. Connect with him on LinkedIn:  Linkedin

Share this post