Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

safe-debug

Rigor Debug / Rigor Audit skill for deep learning research work. Use when the user pastes a traceback, terminal error, CUDA OOM, checkpoint load failure, shape mismatch, NaN loss symptom, or training failure and wants conservative diagnosis before any patching, with debug fixes clearly separated from research contributions. Do not use for broad refactoring, speculative adaptation, automatic exploratory patching, or general repository familiarization.

503 repo starsObserved in 3 repos
Testing & Quality

sql-optimization-patterns

Master SQL query optimization, indexing strategies, and EXPLAIN analysis to dramatically improve database performance and eliminate slow queries. Use when debugging slow queries, designing database schemas, or optimizing application performance.

503 repo starsObserved in 3 repos
Testing & Quality

explore-run

Rigor Improve / Rigor Explore run leaf skill for bounded exploratory evidence in deep learning research repositories. Use when the researcher explicitly authorizes exploratory runs such as small-subset validation, short-cycle guess-and-check, batch sweeps, idle-GPU search, or quick transfer-learning trials, with fair-comparison caveats and no-overclaim summaries in `explore_outputs/`. Do not use for end-to-end exploration orchestration on top of `current_research`, trusted baseline execution, conservative training verification, default routing, verified SOTA claims, or implicit experimentation.

503 repo starsObserved in 2 repos
Testing & Quality

run-train

Rigor Train skill for deep learning research repositories. Use when a documented or selected training command should be run conservatively for startup verification, short-run verification, full kickoff, or resume, with command, config, seed, log, checkpoint, status, and metric evidence written to standardized `train_outputs/`. Do not use for environment setup, exploratory sweeps, speculative idea implementation, or end-to-end orchestration.

503 repo starsObserved in 2 repos
Testing & Quality

minimal-run-and-audit

Rigor Run skill for README-first deep learning repo reproduction. Use when the task is specifically to capture or normalize evidence from the selected smoke test or documented inference or evaluation command and write standardized `repro_outputs/` files, including patch notes when repository files changed. Do not use for training execution, initial repo intake, generic environment setup, paper lookup, target selection, hidden scientific-meaning changes, or end-to-end orchestration by itself.

503 repo starsObserved in 1 repos
Testing & Quality

test-downstream

Test a downstream project against scikit-build-core to check whether the current checkout regresses it versus a released baseline. Builds the project twice per mode — once against a released tag (baseline, via a git worktree) and once against the working tree — in both a normal wheel build and an editable install, then reports parity, wheel-content, and warning diffs. Use this whenever the user wants to test/build/try a downstream or real-world project against their changes, check that a fix or PR doesn't break downstream, run `nox -s downstream`, compare current-vs-released behavior, reproduce a downstream build issue, or validate against projects in docs/data/projects.toml — even if they only name a project (e.g. "test iminuit", "does gemmi still build").

503 repo starsObserved in 1 repos
Testing & Quality

config-loader-helper

Diagnose configuration-related failures, enumerate required env vars, and guide safe local test setup (no secrets).

502 repo starsObserved in 1 repos
Testing & Quality

db-infra-mocks

Propose minimal seams and local substitutes so tests run without real RDBMS/Redis/Mongo infrastructure.

502 repo starsObserved in 1 repos
Testing & Quality

fix-suggester

Diagnose failures and propose minimal, test-backed fixes with verification and rollback instructions.

502 repo starsObserved in 1 repos
Testing & Quality

linter-runner

Execute repository linters and surface highest-priority issues with minimal, targeted fixes.

502 repo starsObserved in 1 repos
Testing & Quality