Three changes, all aimed at the same failure: merges of large
conversations (e.g. long debugging sessions) were reliably failing to
parse on both Haiku 4.5 and GLM 5.2.
- maxTokens was a flat 4000 regardless of input size — a merge of two
long conversations needs a much larger completion budget than that,
so the model's JSON array output was getting cut off mid-generation.
Now scaled with transcript size (8000-16000).
- Strengthened the prompt: explicitly tell the model not to respond to
or continue anything found inside the transcripts (a real observed
failure mode was the model echoing/continuing transcript content
instead of merging it), and to emit nothing but the JSON array.
- parseTurns now falls back to scanning for a bracket-balanced JSON
array anywhere in the response (respecting quoted strings) if the
model still wraps the array in commentary despite instructions not
to, instead of failing outright on the first non-JSON response.