Implements PROMPT.md / docs/REVIEW_PIPELINE.md Phase 2:
- internal/llm/ollama.go — real Ollama provider:
- HealthCheck probes /api/tags + a 1-token completion + a JSON-mode
probe ({"ok": true} round-trip), populating the model-doctor.json
schema documented in docs/LOCAL_MODEL_SETUP.md
- Complete + CompleteJSON via /api/chat with stream=false
- think=false set for ALL completions (qwen3.5:latest is reasoning-
capable but the inner-loop hot path wants direct answers, not
reasoning traces consuming the token budget — same finding as
the Lakehouse-Go chatd 2026-04-30 wave)
- internal/llm/review.go — Reviewer wrapper:
- 2-attempt flow: prompt → parse → repair-prompt → parse
- Strict JSON shape enforced; markdown fences stripped before parse
- Severity normalized to enum; out-of-range confidence clamped
- Per-file chunking (file-level for v0; function-level Phase D+)
- Bounded by review-profile max_file_bytes + max_llm_chunk_chars
- pipeline.go — Phase 2 wired between static scan + report gen:
- --enable-llm flag opts in (off by default — static-only is
cheaper and faster)
- Raw output ALWAYS saved to llm-findings.raw.json (forensics)
- Normalized findings → llm-findings.normalized.json
- LLM findings merged into the report findings list (sourced
"llm" so consumers can filter)
- Receipts honestly mark phase status: "ok" | "degraded" | "skipped"
- cli model doctor — real probes replace the Phase A stub.
Verified:
- model doctor: status="ok" with qwen3.5:latest + qwen3:latest both
loaded, basic_prompt_ok=true, json_mode_ok=true
- insecure-repo with --enable-llm: 9 LLM findings; qwen3.5 correctly
flagged SQLi, RCE, hardcoded credentials as critical with verbatim
evidence; 27s wall for 3 chunks
- clean-repo with --enable-llm: 0 LLM findings, 4 parsed chunks, 2.8s
- self-review with --enable-llm: 77 LLM findings + 83 static; 3 of
~30 chunks needed retry (PROMPT.md, REPORT_SCHEMA.md,
SCRUM_TEST_TEMPLATE.md — all eventually parsed); 5min wall
go vet + go test -short clean. Fixture stray.go now `package fixture`
so go-tooling doesn't choke on the orphan.
Phase D (validator cross-check) + Phase E (memory + diff/rules
subcommands) remain.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
83 lines
1.8 KiB
JSON
83 lines
1.8 KiB
JSON
{
|
|
"run_id": "20260430T060713-f513f6dc",
|
|
"repo_path": "tests/fixtures/insecure-repo",
|
|
"started_at": "2026-04-30T06:07:13.917781613Z",
|
|
"finished_at": "2026-04-30T06:07:13.953011207Z",
|
|
"phases": [
|
|
{
|
|
"name": "repo_intake",
|
|
"status": "ok",
|
|
"output_hash": "540f222456204a27",
|
|
"output_files": [
|
|
"repo-intake.json"
|
|
]
|
|
},
|
|
{
|
|
"name": "static_scan",
|
|
"status": "ok",
|
|
"output_hash": "a7aeccbda6841c1e",
|
|
"output_files": [
|
|
"static-findings.json"
|
|
]
|
|
},
|
|
{
|
|
"name": "llm_review",
|
|
"status": "skipped",
|
|
"errors": [
|
|
"LLM review not requested (pass --enable-llm to opt in)"
|
|
]
|
|
},
|
|
{
|
|
"name": "validation",
|
|
"status": "skipped",
|
|
"errors": [
|
|
"Phase D not implemented in MVP — depends on Phase C"
|
|
]
|
|
},
|
|
{
|
|
"name": "report_generation",
|
|
"status": "ok",
|
|
"output_files": [
|
|
"scrum-test.md",
|
|
"risk-register.md",
|
|
"claim-coverage-table.md",
|
|
"sprint-backlog.md",
|
|
"acceptance-gates.md"
|
|
]
|
|
},
|
|
{
|
|
"name": "memory_update",
|
|
"status": "skipped",
|
|
"errors": [
|
|
"Phase E not implemented in MVP"
|
|
]
|
|
}
|
|
],
|
|
"summary": {
|
|
"total": 16,
|
|
"confirmed": 2,
|
|
"suspected": 14,
|
|
"rejected": 0,
|
|
"critical": 3,
|
|
"high": 4,
|
|
"medium": 6,
|
|
"low": 3,
|
|
"by_source": {
|
|
"static": 16
|
|
},
|
|
"by_check": {
|
|
"static.broad_cors": 1,
|
|
"static.env_file_committed": 1,
|
|
"static.exposed_mutation_endpoint": 2,
|
|
"static.hardcoded_local_ip": 1,
|
|
"static.hardcoded_paths": 1,
|
|
"static.large_files": 1,
|
|
"static.missing_tests": 1,
|
|
"static.raw_sql_interpolation": 1,
|
|
"static.secret_patterns": 3,
|
|
"static.shell_execution": 1,
|
|
"static.todo_comments": 3
|
|
}
|
|
}
|
|
}
|