local-review-harness

profit/local-review-harness

Fork 0

Commit Graph

Author	SHA1	Message	Date
Claude (review-harness setup)	305f5b9ac0	Apply B1+B2+B3+B4 from 2026-04-30 cross-lineage scrum Cross-lineage scrum (Opus 4.7 / Kimi K2.6 / Qwen3-coder via chatd's /v1/chat) on the harness's first 4 commits surfaced 5 real bugs; this commit lands the 4 inside the LLM/validator stack. B5 (scanner skip-list semantics) ships separately as it changes scan behavior on every target repo. B1 (Kimi BLOCK + Opus WARN convergent) — internal/validators: evidencePresent had two flaws: (1) cursor advanced on match in the trim-line fallback, breaking same-line repeated matches AND skipping not-yet-considered lines so out-of-order evidence spuriously failed; (2) strings.Contains on a single `}` trim-matched any closing brace in the file, defeating the "evidence quotes real text" contract. Fix: trivial-evidence guard FIRST (reject anything <4 non-whitespace chars) + per-line search no longer advances a cursor. New regression test TestEvidencePresent_RejectsTrivialMatches covers `}`, `{`, `)`, empty, and out-of-order multi-line evidence (which now passes — order isn't part of the contract). B2 (Kimi WARN + Opus WARN convergent) — internal/pipeline: WriteJSON error for rejected-findings.json was swallowed with `if err == nil`, so a write failure left the validation phase reporting status="ok" while the audit trail vanished. Mirror the validated-findings branch: surface the error in validatePhase.Errors + bump status to degraded + ExitCode=66. B3 (Kimi BLOCK + Opus BLOCK convergent) — internal/llm/ollama.go: HealthCheck.basic_prompt_ok was set to true on ANY non-empty response, so a model emitting `<think>...` traces or apologies passed silently. Now requires the response to contain "OK" (uppercase, substring). Substring rather than equality lets minor whitespace/punctuation variations through (some models add a trailing period). Errors now record what the model actually said when it fails the check. B4 (Opus BLOCK only — same class as today's chatd Anthropic-temp fix) — internal/llm/ollama.go: chatBody had `if opts.Temperature != 0` which silently dropped Temperature=0 from the request, so HealthCheck + Reviewer (both pass Temperature=0 expecting determinism) actually ran at Ollama's ~0.8 default. Always forward Temperature now. The two callers always set explicit values, so "0 means 0" is correct; if a future caller wants Ollama's default they'll switch CompleteOptions.Temperature to *float64 like chatd did this morning. Verified end-to-end: insecure-repo + --enable-llm still produces 25 confirmed findings (16 static + 9 LLM), 0 rejected. Validator unit tests: 11 pass (added TestEvidencePresent_RejectsTrivialMatches). Same-day-as-shipping scrum, same-day-as-shipping fixes. The convergent-≥2 gate caught 3 of these; the 4th was Opus-only but verified by reading the code (same idiom as today's chatd bug). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-04-30 01:33:01 -05:00
Claude (review-harness setup)	4dc53c5798	Phase D — validator cross-checks LLM findings + 2 close-out fixes Implements PROMPT.md / docs/REVIEW_PIPELINE.md Phase 3: "AI may suggest. Code validates." internal/validators/validate.go — 3 hard checks per the "Reject A Finding If" list: - file does not exist (with path-traversal guard against the LLM hallucinating ../../../etc/passwd) - cited evidence does not appear in the file (verbatim or trim-line-by-line — models often re-indent quotes when quoting code) - line hint exceeds file line count 3 soft checks documented as open (claim semantics, suggested-fix relevance, invented tests/commands — all need another LLM pass). internal/validators/validate_test.go — 9 tests including: - TestValidate_RejectsNonexistentFile (gate D1) - TestValidate_RejectsEvidenceNotInFile - TestValidate_RejectsLineHintBeyondFile - TestValidate_AcceptsRealFinding - TestValidate_AcceptsEvidenceWithDifferentLeadingWhitespace - TestValidate_RejectsEmptyEvidence - TestValidate_PassesThroughStaticFindings - TestValidate_RejectsPathEscapingRepo (path-traversal protection) - TestValidate_AcceptsRelativeRepoPath (the regression — see below) Pipeline phase 3 wired between LLM review (Phase C) and report gen (Phase 4). validated-findings.json contains the confirmed set; rejected-findings.json contains rejects with per-finding reason + detail. Receipt phase entry honest about output files + status. === Bug J caught === First Phase D run rejected EVERY real LLM finding as file_not_found because the path-traversal check compared a relative joined path (`tests/fixtures/insecure-repo/src/handler.go`) against an absolute repoAbs (`/home/profit/share/.../insecure-repo`), so HasPrefix always returned false. Both sides now resolved via filepath.Abs before comparison. Regression test TestValidate_AcceptsRelativeRepoPath locks this in — runs the validator against a relative repo path AND a relative chdir, the exact shape that hit the bug. J's framing was honest: "I don't know what the problem is, but you know what we're trying to accomplish." The fix-it-yourself signal let me trace through the rejection details + see the smoking gun in the detail string ("escapes repo root"). Without that prompt the 9 false rejections might have looked like real LLM bugs. === 2 close-out fixes === 1. .gitignore: changed `/reports/latest/` → `*/reports/latest/` (and same for `run-`). Phase C committed 22 generated files from `tests/fixtures/*/reports/latest/` because the original pattern was anchored at the harness root only. Existing tracked files removed via git rm --cached; new pattern keeps fixture reports out of version control going forward. 2. pipeline.cleanOutputDir: pipeline now deletes the bounded list of known per-run files at the start of each run. Before this, a prior run's rejected-findings.json could linger when the current run had no rejections — confused J during the bug hunt above. cleanOutputDir is bounded (deletes only files we emit) so operator-owned adjacent files stay. Verified end-to-end: insecure-repo + --enable-llm → 25 confirmed findings (16 static + 9 LLM), 0 rejected. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-04-30 01:24:02 -05:00

Author

SHA1

Message

Date

Claude (review-harness setup)

305f5b9ac0

Apply B1+B2+B3+B4 from 2026-04-30 cross-lineage scrum

Cross-lineage scrum (Opus 4.7 / Kimi K2.6 / Qwen3-coder via chatd's
/v1/chat) on the harness's first 4 commits surfaced 5 real bugs;
this commit lands the 4 inside the LLM/validator stack. B5 (scanner
skip-list semantics) ships separately as it changes scan behavior
on every target repo.

B1 (Kimi BLOCK + Opus WARN convergent) — internal/validators:
evidencePresent had two flaws: (1) cursor advanced on match in the
trim-line fallback, breaking same-line repeated matches AND skipping
not-yet-considered lines so out-of-order evidence spuriously failed;
(2) strings.Contains on a single `}` trim-matched any closing brace
in the file, defeating the "evidence quotes real text" contract.
Fix: trivial-evidence guard FIRST (reject anything <4 non-whitespace
chars) + per-line search no longer advances a cursor. New regression
test TestEvidencePresent_RejectsTrivialMatches covers `}`, `{`, `)`,
empty, and out-of-order multi-line evidence (which now passes —
order isn't part of the contract).

B2 (Kimi WARN + Opus WARN convergent) — internal/pipeline:
WriteJSON error for rejected-findings.json was swallowed with
`if err == nil`, so a write failure left the validation phase
reporting status="ok" while the audit trail vanished. Mirror the
validated-findings branch: surface the error in
validatePhase.Errors + bump status to degraded + ExitCode=66.

B3 (Kimi BLOCK + Opus BLOCK convergent) — internal/llm/ollama.go:
HealthCheck.basic_prompt_ok was set to true on ANY non-empty
response, so a model emitting `<think>...` traces or apologies
passed silently. Now requires the response to contain "OK"
(uppercase, substring). Substring rather than equality lets minor
whitespace/punctuation variations through (some models add a
trailing period). Errors now record what the model actually said
when it fails the check.

B4 (Opus BLOCK only — same class as today's chatd Anthropic-temp
fix) — internal/llm/ollama.go: chatBody had `if opts.Temperature != 0`
which silently dropped Temperature=0 from the request, so HealthCheck
+ Reviewer (both pass Temperature=0 expecting determinism) actually
ran at Ollama's ~0.8 default. Always forward Temperature now. The
two callers always set explicit values, so "0 means 0" is correct;
if a future caller wants Ollama's default they'll switch
CompleteOptions.Temperature to *float64 like chatd did this morning.

Verified end-to-end: insecure-repo + --enable-llm still produces 25
confirmed findings (16 static + 9 LLM), 0 rejected. Validator unit
tests: 11 pass (added TestEvidencePresent_RejectsTrivialMatches).

Same-day-as-shipping scrum, same-day-as-shipping fixes. The
convergent-≥2 gate caught 3 of these; the 4th was Opus-only but
verified by reading the code (same idiom as today's chatd bug).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

2026-04-30 01:33:01 -05:00

Claude (review-harness setup)

4dc53c5798

Phase D — validator cross-checks LLM findings + 2 close-out fixes

Implements PROMPT.md / docs/REVIEW_PIPELINE.md Phase 3:
"AI may suggest. Code validates."

internal/validators/validate.go — 3 hard checks per the
"Reject A Finding If" list:
- file does not exist (with path-traversal guard against the LLM
  hallucinating ../../../etc/passwd)
- cited evidence does not appear in the file (verbatim or
  trim-line-by-line — models often re-indent quotes when quoting code)
- line hint exceeds file line count

3 soft checks documented as open (claim semantics, suggested-fix
relevance, invented tests/commands — all need another LLM pass).

internal/validators/validate_test.go — 9 tests including:
- TestValidate_RejectsNonexistentFile (gate D1)
- TestValidate_RejectsEvidenceNotInFile
- TestValidate_RejectsLineHintBeyondFile
- TestValidate_AcceptsRealFinding
- TestValidate_AcceptsEvidenceWithDifferentLeadingWhitespace
- TestValidate_RejectsEmptyEvidence
- TestValidate_PassesThroughStaticFindings
- TestValidate_RejectsPathEscapingRepo (path-traversal protection)
- TestValidate_AcceptsRelativeRepoPath (the regression — see below)

Pipeline phase 3 wired between LLM review (Phase C) and report gen
(Phase 4). validated-findings.json contains the confirmed set;
rejected-findings.json contains rejects with per-finding reason +
detail. Receipt phase entry honest about output files + status.

=== Bug J caught ===

First Phase D run rejected EVERY real LLM finding as file_not_found
because the path-traversal check compared a relative joined path
(`tests/fixtures/insecure-repo/src/handler.go`) against an absolute
repoAbs (`/home/profit/share/.../insecure-repo`), so HasPrefix
always returned false. Both sides now resolved via filepath.Abs
before comparison. Regression test
TestValidate_AcceptsRelativeRepoPath locks this in — runs the
validator against a relative repo path AND a relative chdir, the
exact shape that hit the bug.

J's framing was honest: "I don't know what the problem is, but you
know what we're trying to accomplish." The fix-it-yourself signal
let me trace through the rejection details + see the smoking gun
in the detail string ("escapes repo root"). Without that prompt the
9 false rejections might have looked like real LLM bugs.

=== 2 close-out fixes ===

1. .gitignore: changed `/reports/latest/` → `**/reports/latest/`
   (and same for `run-*`). Phase C committed 22 generated files
   from `tests/fixtures/*/reports/latest/` because the original
   pattern was anchored at the harness root only. Existing tracked
   files removed via git rm --cached; new pattern keeps fixture
   reports out of version control going forward.

2. pipeline.cleanOutputDir: pipeline now deletes the bounded list
   of known per-run files at the start of each run. Before this,
   a prior run's rejected-findings.json could linger when the
   current run had no rejections — confused J during the bug hunt
   above. cleanOutputDir is bounded (deletes only files we emit)
   so operator-owned adjacent files stay.

Verified end-to-end: insecure-repo + --enable-llm → 25 confirmed
findings (16 static + 9 LLM), 0 rejected.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

2026-04-30 01:24:02 -05:00

2 Commits