Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

vlm-benchmark

Benchmark an optimization (PR or branch) on a Device Farm device via the vlm-benchmark framework — baseline vs optimized, quality-regression-aware.

315 repo starsObserved in 1 repos
Testing & Quality

bisq-pr-reviewer

Provide comprehensive Bisq PR review in Codex or Claude Code by integrating CodeRabbitAI feedback extraction with Bisq-specific security validation, contribution standards verification, and domain expertise for Bitcoin, P2P, trade protocol, DAO, Java, and JavaFX changes. Use when reviewing Bisq pull requests, validating contribution standards, extracting unresolved review comments, or performing security reviews of cryptocurrency code changes.

314 repo starsObserved in 1 repos
Testing & Quality

perf-experiment

Run WebAssembly-in-Chrome performance experiments for Mandelbrot tile generation - build wasm variants with alternative compiler/optimizer flags, benchmark them against a baseline across real and synthetic workloads, verify correctness, and apply measured winners to the production config. Use when asked to optimize tile generation, wasm performance, or run/interpret benchmarks.

314 repo starsObserved in 1 repos
Testing & Quality

systematic-debug

系统化调试技能 - 按证据驱动流程定位和修复bug

314 repo starsObserved in 1 repos
Testing & Quality

prism-full

Full Prism: multi-pass structural analysis with mandatory adversarial self-correction. Designs custom analytical passes, executes them with chaining, then attacks its own findings before synthesizing. Use for maximum depth on important code or artifacts.

313 repo starsObserved in 5 repos
Testing & Quality

prism-scan

Structural analysis through dynamically generated cognitive lenses. Generates the optimal analytical lens for the specific code/artifact, then executes it. Finds conservation laws, structural invariants, and concrete bugs that vanilla analysis misses. Use on any code file, system design, or text artifact.

313 repo starsObserved in 4 repos
Testing & Quality

vitest-evals

Use when authoring, reviewing, or debugging harness-backed vitest-evals suites, custom Harness adapters, first-party ai-sdk or pi-ai harness integrations, judges, replay, reporter-facing normalized run data, or examples and docs for these APIs.

313 repo starsObserved in 1 repos
Testing & Quality

packet-validation

Add or run libcrafter packet behavior validation through oracle specs, backend adapters, offline checks, pcap checks, live checks, and artifacts. Use when packet behavior changes need validation coverage.

312 repo starsObserved in 1 repos
Testing & Quality

deep-review-task

Run the sandboxed deep-review tool against a benchmark task PR — launches Claude Code in Docker with pre-fetched PR artifacts and writes review-summary.md / issues-found.md

311 repo starsObserved in 2 repos
Testing & Quality

fin-guru-compliance-review

Execute comprehensive compliance reviews for Finance Guru deliverables. Validates disclaimers, data handling, risk disclosures, and regulatory positioning.

311 repo starsObserved in 1 repos
Testing & Quality