local-review-harness

profit/local-review-harness

Fork 0

Commit Graph

Author	SHA1	Message	Date
Claude (review-harness setup)	305f5b9ac0	Apply B1+B2+B3+B4 from 2026-04-30 cross-lineage scrum Cross-lineage scrum (Opus 4.7 / Kimi K2.6 / Qwen3-coder via chatd's /v1/chat) on the harness's first 4 commits surfaced 5 real bugs; this commit lands the 4 inside the LLM/validator stack. B5 (scanner skip-list semantics) ships separately as it changes scan behavior on every target repo. B1 (Kimi BLOCK + Opus WARN convergent) — internal/validators: evidencePresent had two flaws: (1) cursor advanced on match in the trim-line fallback, breaking same-line repeated matches AND skipping not-yet-considered lines so out-of-order evidence spuriously failed; (2) strings.Contains on a single `}` trim-matched any closing brace in the file, defeating the "evidence quotes real text" contract. Fix: trivial-evidence guard FIRST (reject anything <4 non-whitespace chars) + per-line search no longer advances a cursor. New regression test TestEvidencePresent_RejectsTrivialMatches covers `}`, `{`, `)`, empty, and out-of-order multi-line evidence (which now passes — order isn't part of the contract). B2 (Kimi WARN + Opus WARN convergent) — internal/pipeline: WriteJSON error for rejected-findings.json was swallowed with `if err == nil`, so a write failure left the validation phase reporting status="ok" while the audit trail vanished. Mirror the validated-findings branch: surface the error in validatePhase.Errors + bump status to degraded + ExitCode=66. B3 (Kimi BLOCK + Opus BLOCK convergent) — internal/llm/ollama.go: HealthCheck.basic_prompt_ok was set to true on ANY non-empty response, so a model emitting `<think>...` traces or apologies passed silently. Now requires the response to contain "OK" (uppercase, substring). Substring rather than equality lets minor whitespace/punctuation variations through (some models add a trailing period). Errors now record what the model actually said when it fails the check. B4 (Opus BLOCK only — same class as today's chatd Anthropic-temp fix) — internal/llm/ollama.go: chatBody had `if opts.Temperature != 0` which silently dropped Temperature=0 from the request, so HealthCheck + Reviewer (both pass Temperature=0 expecting determinism) actually ran at Ollama's ~0.8 default. Always forward Temperature now. The two callers always set explicit values, so "0 means 0" is correct; if a future caller wants Ollama's default they'll switch CompleteOptions.Temperature to *float64 like chatd did this morning. Verified end-to-end: insecure-repo + --enable-llm still produces 25 confirmed findings (16 static + 9 LLM), 0 rejected. Validator unit tests: 11 pass (added TestEvidencePresent_RejectsTrivialMatches). Same-day-as-shipping scrum, same-day-as-shipping fixes. The convergent-≥2 gate caught 3 of these; the 4th was Opus-only but verified by reading the code (same idiom as today's chatd bug). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-04-30 01:33:01 -05:00
Claude (review-harness setup)	e346b54e0f	Phase C — local-Ollama LLM review wired end-to-end Implements PROMPT.md / docs/REVIEW_PIPELINE.md Phase 2: - internal/llm/ollama.go — real Ollama provider: - HealthCheck probes /api/tags + a 1-token completion + a JSON-mode probe ({"ok": true} round-trip), populating the model-doctor.json schema documented in docs/LOCAL_MODEL_SETUP.md - Complete + CompleteJSON via /api/chat with stream=false - think=false set for ALL completions (qwen3.5:latest is reasoning- capable but the inner-loop hot path wants direct answers, not reasoning traces consuming the token budget — same finding as the Lakehouse-Go chatd 2026-04-30 wave) - internal/llm/review.go — Reviewer wrapper: - 2-attempt flow: prompt → parse → repair-prompt → parse - Strict JSON shape enforced; markdown fences stripped before parse - Severity normalized to enum; out-of-range confidence clamped - Per-file chunking (file-level for v0; function-level Phase D+) - Bounded by review-profile max_file_bytes + max_llm_chunk_chars - pipeline.go — Phase 2 wired between static scan + report gen: - --enable-llm flag opts in (off by default — static-only is cheaper and faster) - Raw output ALWAYS saved to llm-findings.raw.json (forensics) - Normalized findings → llm-findings.normalized.json - LLM findings merged into the report findings list (sourced "llm" so consumers can filter) - Receipts honestly mark phase status: "ok" \| "degraded" \| "skipped" - cli model doctor — real probes replace the Phase A stub. Verified: - model doctor: status="ok" with qwen3.5:latest + qwen3:latest both loaded, basic_prompt_ok=true, json_mode_ok=true - insecure-repo with --enable-llm: 9 LLM findings; qwen3.5 correctly flagged SQLi, RCE, hardcoded credentials as critical with verbatim evidence; 27s wall for 3 chunks - clean-repo with --enable-llm: 0 LLM findings, 4 parsed chunks, 2.8s - self-review with --enable-llm: 77 LLM findings + 83 static; 3 of ~30 chunks needed retry (PROMPT.md, REPORT_SCHEMA.md, SCRUM_TEST_TEMPLATE.md — all eventually parsed); 5min wall go vet + go test -short clean. Fixture stray.go now `package fixture` so go-tooling doesn't choke on the orphan. Phase D (validator cross-check) + Phase E (memory + diff/rules subcommands) remain. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-04-30 01:13:39 -05:00
Claude (review-harness setup)	f3ee4722a8	Phase A + B (MVP) — local review harness Implements the MVP cutline from the planning artifact: - Phase A: skeleton + CLI dispatch + provider interface + stub model doctor - Phase B: scanner + git probe + 12 static analyzers + reporters + pipeline - Phase B fixtures: clean-repo, insecure-repo, degraded-repo 12 static analyzers per PROMPT.md "Suggested Static Checks For MVP": hardcoded_paths, shell_execution, raw_sql_interpolation, broad_cors, secret_patterns, large_files, todo_comments, missing_tests, env_file_committed, unsafe_file_io, exposed_mutation_endpoint, hardcoded_local_ip. Acceptance gates passing: - B1 (intake produces accurate counts) ✓ - B2 (insecure fixture fires ≥8 distinct check_ids — actually 11/12) ✓ - B3 (clean fixture produces 0 confirmed findings — no false positives) ✓ - B4 (scrum mode produces all 6 required markdown + JSON reports) ✓ - B5 (receipts.json marks degraded phases honestly) ✓ - F (self-review on this repo runs without crashing) ✓ — exit 66 (degraded because Phase C LLM review is hardcoded skipped) Phases C (LLM review), D (validation cross-check), E (memory + diff + rules subcommands) deferred per the cutline. The MVP delivers the evidence-first path; LLM is purely additive. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>	2026-04-30 00:56:02 -05:00

Author

SHA1

Message

Date

Claude (review-harness setup)

305f5b9ac0

Apply B1+B2+B3+B4 from 2026-04-30 cross-lineage scrum

Cross-lineage scrum (Opus 4.7 / Kimi K2.6 / Qwen3-coder via chatd's
/v1/chat) on the harness's first 4 commits surfaced 5 real bugs;
this commit lands the 4 inside the LLM/validator stack. B5 (scanner
skip-list semantics) ships separately as it changes scan behavior
on every target repo.

B1 (Kimi BLOCK + Opus WARN convergent) — internal/validators:
evidencePresent had two flaws: (1) cursor advanced on match in the
trim-line fallback, breaking same-line repeated matches AND skipping
not-yet-considered lines so out-of-order evidence spuriously failed;
(2) strings.Contains on a single `}` trim-matched any closing brace
in the file, defeating the "evidence quotes real text" contract.
Fix: trivial-evidence guard FIRST (reject anything <4 non-whitespace
chars) + per-line search no longer advances a cursor. New regression
test TestEvidencePresent_RejectsTrivialMatches covers `}`, `{`, `)`,
empty, and out-of-order multi-line evidence (which now passes —
order isn't part of the contract).

B2 (Kimi WARN + Opus WARN convergent) — internal/pipeline:
WriteJSON error for rejected-findings.json was swallowed with
`if err == nil`, so a write failure left the validation phase
reporting status="ok" while the audit trail vanished. Mirror the
validated-findings branch: surface the error in
validatePhase.Errors + bump status to degraded + ExitCode=66.

B3 (Kimi BLOCK + Opus BLOCK convergent) — internal/llm/ollama.go:
HealthCheck.basic_prompt_ok was set to true on ANY non-empty
response, so a model emitting `<think>...` traces or apologies
passed silently. Now requires the response to contain "OK"
(uppercase, substring). Substring rather than equality lets minor
whitespace/punctuation variations through (some models add a
trailing period). Errors now record what the model actually said
when it fails the check.

B4 (Opus BLOCK only — same class as today's chatd Anthropic-temp
fix) — internal/llm/ollama.go: chatBody had `if opts.Temperature != 0`
which silently dropped Temperature=0 from the request, so HealthCheck
+ Reviewer (both pass Temperature=0 expecting determinism) actually
ran at Ollama's ~0.8 default. Always forward Temperature now. The
two callers always set explicit values, so "0 means 0" is correct;
if a future caller wants Ollama's default they'll switch
CompleteOptions.Temperature to *float64 like chatd did this morning.

Verified end-to-end: insecure-repo + --enable-llm still produces 25
confirmed findings (16 static + 9 LLM), 0 rejected. Validator unit
tests: 11 pass (added TestEvidencePresent_RejectsTrivialMatches).

Same-day-as-shipping scrum, same-day-as-shipping fixes. The
convergent-≥2 gate caught 3 of these; the 4th was Opus-only but
verified by reading the code (same idiom as today's chatd bug).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

2026-04-30 01:33:01 -05:00

Claude (review-harness setup)

e346b54e0f

Phase C — local-Ollama LLM review wired end-to-end

Implements PROMPT.md / docs/REVIEW_PIPELINE.md Phase 2:
- internal/llm/ollama.go — real Ollama provider:
  - HealthCheck probes /api/tags + a 1-token completion + a JSON-mode
    probe ({"ok": true} round-trip), populating the model-doctor.json
    schema documented in docs/LOCAL_MODEL_SETUP.md
  - Complete + CompleteJSON via /api/chat with stream=false
  - think=false set for ALL completions (qwen3.5:latest is reasoning-
    capable but the inner-loop hot path wants direct answers, not
    reasoning traces consuming the token budget — same finding as
    the Lakehouse-Go chatd 2026-04-30 wave)
- internal/llm/review.go — Reviewer wrapper:
  - 2-attempt flow: prompt → parse → repair-prompt → parse
  - Strict JSON shape enforced; markdown fences stripped before parse
  - Severity normalized to enum; out-of-range confidence clamped
  - Per-file chunking (file-level for v0; function-level Phase D+)
  - Bounded by review-profile max_file_bytes + max_llm_chunk_chars
- pipeline.go — Phase 2 wired between static scan + report gen:
  - --enable-llm flag opts in (off by default — static-only is
    cheaper and faster)
  - Raw output ALWAYS saved to llm-findings.raw.json (forensics)
  - Normalized findings → llm-findings.normalized.json
  - LLM findings merged into the report findings list (sourced
    "llm" so consumers can filter)
  - Receipts honestly mark phase status: "ok" | "degraded" | "skipped"
- cli model doctor — real probes replace the Phase A stub.

Verified:
- model doctor: status="ok" with qwen3.5:latest + qwen3:latest both
  loaded, basic_prompt_ok=true, json_mode_ok=true
- insecure-repo with --enable-llm: 9 LLM findings; qwen3.5 correctly
  flagged SQLi, RCE, hardcoded credentials as critical with verbatim
  evidence; 27s wall for 3 chunks
- clean-repo with --enable-llm: 0 LLM findings, 4 parsed chunks, 2.8s
- self-review with --enable-llm: 77 LLM findings + 83 static; 3 of
  ~30 chunks needed retry (PROMPT.md, REPORT_SCHEMA.md,
  SCRUM_TEST_TEMPLATE.md — all eventually parsed); 5min wall

go vet + go test -short clean. Fixture stray.go now `package fixture`
so go-tooling doesn't choke on the orphan.

Phase D (validator cross-check) + Phase E (memory + diff/rules
subcommands) remain.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

2026-04-30 01:13:39 -05:00

Claude (review-harness setup)

f3ee4722a8

Phase A + B (MVP) — local review harness

Implements the MVP cutline from the planning artifact:
- Phase A: skeleton + CLI dispatch + provider interface + stub model doctor
- Phase B: scanner + git probe + 12 static analyzers + reporters + pipeline
- Phase B fixtures: clean-repo, insecure-repo, degraded-repo

12 static analyzers per PROMPT.md "Suggested Static Checks For MVP":
hardcoded_paths, shell_execution, raw_sql_interpolation, broad_cors,
secret_patterns, large_files, todo_comments, missing_tests,
env_file_committed, unsafe_file_io, exposed_mutation_endpoint,
hardcoded_local_ip.

Acceptance gates passing:
- B1 (intake produces accurate counts) ✓
- B2 (insecure fixture fires ≥8 distinct check_ids — actually 11/12) ✓
- B3 (clean fixture produces 0 confirmed findings — no false positives) ✓
- B4 (scrum mode produces all 6 required markdown + JSON reports) ✓
- B5 (receipts.json marks degraded phases honestly) ✓
- F  (self-review on this repo runs without crashing) ✓ — exit 66 (degraded
  because Phase C LLM review is hardcoded skipped)

Phases C (LLM review), D (validation cross-check), E (memory + diff +
rules subcommands) deferred per the cutline. The MVP delivers the
evidence-first path; LLM is purely additive.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

2026-04-30 00:56:02 -05:00

3 Commits