Skip to content

fix(ce-noslop): keep claim certainty and rule out semicolons as dash substitutes - #1847

Merged
tmchow merged 2 commits into
mainfrom
t3/review-humanizer-skill
Oct 10, 2026
Merged

tmchow merged 2 commits into
mainfrom
t3/review-humanizer-skill

Conversation

@tmchow

@tmchow tmchow commented Oct 10, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

ce-noslop now treats a rewrite that changes how certain a claim is as a failure, not only one that drops a qualifier, and its em dash rule no longer allows a semicolon as the replacement. Both came from reviewing phuryn/work-humanizer for ideas. A third idea from that review was tried and removed because evals showed it made the skill drop claims.

What was tried and removed

A "What a fix may do" paragraph in references/patterns.md said that fixes must keep certainty, must cut a claim (edit mode) or name the missing fact (detect mode) when a specific is missing, and must not leave clipped sentences or add personality. On longer drafts it made things worse:

Check (longer drafts, 2 trials per host) Before With the paragraph
Release note keeps "We think this is the most significant improvement since 3.0" Claude 2/2, Codex 2/2 Claude 0/2, Codex 2/2
Status update keeps "cautiously optimistic, but not done" 3/4 1/4
Fragment stack left in the result 0/4 1/4

The "cut the claim" instruction fired on hedged opinions as well as puffery. The clipped-sentence guard did not prevent the one fragment stack. Detect mode already named missing evidence without inventing it (4/4 before the change).

Validation

  • Short fixtures, before and after, on Claude and Codex: no difference, because the baseline already behaved correctly.
  • Longer fixtures (release note, status update, support reply, blog post, plus detect on the blog post), 40 runs: the results above. Before the change, Codex once turned "remove most tickets" into "eliminate most tickets" and once replaced a dash with a semicolon. Neither happened after the change. That is one instance each, so it is weak evidence for the two kept lines.
  • Four new eval-cell rows with fixtures: fixes-leave-full-sentences, edit-keeps-claim-certainty, long-release-note-keeps-hedged-opinion, long-status-update-keeps-hedges, baselined at 67035e9. All 8 post-change cells pass on Claude and Codex with the paragraph removed.
  • The rows check the rewritten text itself through a new result_must_include grade, which reads only the RESULT block. That way a summary line that names a kept hedge can't satisfy the check. The semicolon row is post-only, since the baseline rule 29 did not prohibit semicolons.
  • bun run test, bun run typecheck, and bun run release:validate pass.

Security Disclosure

No security-relevant changes.

Agent Disclosure

  • Model: Claude Code · claude-opus-5-5

…substitutes

A rewrite now fails when it changes how certain a claim is, not only
when it drops a qualifier. The em dash rule says a semicolon is not a
substitute. Four eval scenarios cover the dash rewrite, hedged claims,
and longer drafts where an opinion and real hedges must survive.
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-10T08:13:55.330297Z 20ab2a4 New commits
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0cf3bea58b

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/skill-eval-cell/catalog.ts Outdated
Comment thread tests/skill-eval-cell/catalog.ts Outdated
The certainty rows passed when a rewrite dropped every hedge but kept
the numbers, because must_include also reads the summary line. A new
result_must_include grade reads only the RESULT block, and the rows now
require the hedged wording itself. The semicolon row is post-only, since
the baseline rule 29 did not prohibit semicolons.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 20ab2a41af

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread tests/skill-eval-cell/catalog.ts
@tmchow
tmchow merged commit 51aa537 into main Oct 10, 2026
9 checks passed
@github-actions github-actions Bot mentioned this pull request Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant