Skip to content

fix(eval): require explicit reasoning_effort for GLM-5.3+ - #59

Merged
ronaldtse merged 1 commit into
mainfrom
fix/glm-reasoning-effort
Aug 30, 2026
Merged

ronaldtse merged 1 commit into
mainfrom
fix/glm-reasoning-effort

Conversation

@ronaldtse

Copy link
Copy Markdown
Contributor

Why

GLM-5.3-Flash silently defaults reasoning_effort to MAX when the
parameter is absent or unrecognized — and thinking: {"type": "disabled"} (the GLM-4.x/5.2 mechanism our eval uses) is exactly an
unrecognized value there. Our 2026-08-17 GLM-5.2 reproduction
recorded reasoning mode at minutes per long paragraph; on the
1,200-paragraph SadeedDiac-25 sweep a missing parameter would turn a
~1h run into days — silently, with no error to catch.

What changes

eval_sadeed_glm.py (the only LLM API call site across rababa /
ml-models / secryst / api — audited):

  • glm-5.3* models refuse to start without an explicit effort
    value: python eval_sadeed_glm.py glm-5.3-flash low
  • When given, both knobs are sent (reasoning_effort +
    thinking.type=disabled) so old and new models are covered
  • Responses containing reasoning_content print a WARNING — a
    tripwire proving the disable did not take
  • The effort level is part of the checkpoint filename (different
    effort = different protocol = separate resume state) and the
    startup protocol line, keeping decode-protocol disclosure intact

Verification

  • py_compile clean; refusal path exercised:
    python eval_sadeed_glm.py glm-5.3-flash → exits with the message
  • glm-5.2 default path unchanged (no effort → payload identical to
    the Aug 17 reproduction that scored 2.5060 raw / 2.6911 zero-skip)

Notes

  • Valid effort values for the 5.3 API should be confirmed against
    z.ai docs on the first real run; an invalid value also falls back
    to MAX, which is why the response tripwire matters.
  • The GLM-5.3 attempt on 2026-08-17 failed with HTTP 403 (key denied
    access) — retrying the leaderboard row is a separate decision.

GLM-5.3-Flash silently defaults absent or unrecognized
reasoning_effort to MAX; thinking.type=disabled is ignored there.
The 2026-08-17 run measured reasoning mode at minutes per long
paragraph, so a missing parameter turns a ~1h SadeedDiac-25 sweep
into days. glm-5.3* models now refuse to start without an effort
value; when given, both knobs are sent and reasoning_content in any
response warns that the disable did not take. Checkpoints and the
startup protocol line carry the effort level.
@ronaldtse
ronaldtse merged commit ea23927 into main Aug 30, 2026
4 of 10 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant