multi-pr-review
Run a consensus-style multi-agent review of a PR with severity-based findings.
Browse reusable Agent Skills, each with a clear purpose and practical guidance.
Run a consensus-style multi-agent review of a PR with severity-based findings.
Scan open pull requests, assess risk and complexity, and build a prioritized review queue. Used by digital twin personas to reduce PR review cognitive load.
Automated visual QA testing using Playwright — navigate web apps like a real user, capture screenshots, find bugs, and fix them.
Map system architecture to ablatable units for ablation studies
Design ablation studies to isolate component contributions in ML systems
Tactic: Construct detailed hostile persona, attack artifact from that persona's perspective, record successful attack paths for aggregation.
Campaign: Logical extreme and boundary testing via reductio ad absurdum and edge-case analysis. Core question: Does this artifact collapse under logical limits and boundary conditions? Methods: Lakatos 1976, Dutilh Novaes 2016, BVA, Flyvbjerg Critical Case, Popper.
SOP: Run the external ARA rigor-reviewer (Seal Level 2, six-dimension semantic review) over ../ara/ and pass its level2_report.json to the user
Detect annotation artifacts and shortcuts in benchmarks
SOP for validating that candidate axes are independent and meaningful.