PaperTool

Author	SHA1	Message	Date
hc	6b78dc47fa	style(agents): standardize bilingual format for all agent files - Use English for structural headers (Role, Workflow, Constraints) - Use Chinese for business logic and detailed explanations - Consistent formatting across all 6 agents: - paper-director.md - paper-analyzer.md - paper-image-extractor.md - code-writer.md - test-runner.md - result-verifier.md	2026-04-01 00:42:01 +08:00
hc	ced50ea2b0	feat(agent): add result-verifier for blind visual comparison Root cause: test-runner was giving overly optimistic results due to: 1. Context bias - knew the implementation, tended to defend it 2. No actual visual comparison - just wrote 'ACCEPTABLE' without looking 3. No structural validation - accepted 35x scale differences as 'acceptable' Solution: - New result-verifier agent that performs blind visual comparison - Strict pass/fail criteria for structural validation - Updated test-runner to use result-verifier for each figure - Clear guidelines: structural mismatches = FAIL, not ACCEPTABLE Test result: verifier correctly identified Fig3 as FAIL with 7 specific issues: - Wrong X-axis variable (channels vs power) - Wrong Y-axis scale (5x difference) - Wrong curve count (5 vs 4) - etc.	2026-03-31 23:56:36 +08:00
hc	5d5aee1f83	refactor: improve verification workflow with visual comparison Major changes: - paper-image-extractor: Generate reference_plots.py for visual verification - paper-director: Add image understanding checkpoint with side-by-side comparison - paper-analyzer: Add data source labeling with reliability levels - code-writer: Change from TDD to VDD (Verification-Driven Development) - test-runner: Generate comparison reports with images and explanations - verification skill: Add difference classification system - code-generation skill: Emphasize result independence Key principles: - Code results are authoritative, paper values are references - Differences are expected and documented, not bugs to fix - Visual comparison prioritized over exact numerical match - Tests verify sanity (shape, gradient, range), not exact values	2026-03-31 19:55:36 +08:00
hc	db731f6745	fix(agents): remove invalid 'model: inherit' configuration OpenCode requires models to be either explicitly defined with valid IDs or omitted to inherit the default model.	2026-03-31 18:08:10 +08:00
hc	f62129f5d4	feat(agents): add test-runner subagent	2026-03-31 17:36:53 +08:00

5 Commits