Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

debug-physics

Diagnose physics misbehaviour in the Genie Sim RT Engine — robot swings on spawn, contacts tunnel, joints drift past their limits, cloth blows up, the convex-hull proxy renders instead of the visual mesh, or the wrong physics backend is active. Walks the user through the engine's debug toggles (visualizers, marker array, GL viewer, `init_*` teleport, backend swap) and the common failure-mode fixes. Trigger: When the user reports "robot swings at start", "objects float / sink into the floor", "contact tunnelling", "joint went past limit", "robot vibrates / explodes", "shelf looks like a convex hull", "wrong gripper poses", "newton vs physx vs mjwarp difference", or asks to "debug the physics".

1.14k repo starsObserved in 2 repos
Testing & Quality

run-benchmark

Launch a geniesim_benchmark task locally (typically inside the GUI Docker container) against a user-provided inference server, using the `geniesim benchmark run` CLI verb. Trigger: When the user asks to "run geniesim", "本地跑仿真", "启动仿真任务", "run a benchmark", "launch <some>_<config>.yaml", or wants to execute a benchmark task config (anything under `geniesim_benchmark/config/*.yaml`) against a remote inference host (ip:port).

1.14k repo starsObserved in 2 repos
Testing & Quality

rocq-build-troubleshoot

Fast workflow to diagnose and fix Rocq/Coq compile errors in this repository, especially missing imports after links/simulate splits and per-file compile checks.

1.14k repo starsObserved in 1 repos
Testing & Quality

llava-onevision2-consistency

Bilingual guide for running and interpreting LLaVA-OneVision2 HF vs Megatron consistency checks across TP and PP settings

1.14k repo starsObserved in 1 repos
Testing & Quality

electrobun-cdp-debug

Debug the LLM Space desktop Electrobun renderer through CEF Chrome DevTools Protocol. Use the bundled raw CDP probe for the actual Electrobun CEF renderer, and use chrome-devtools-axi for ordinary Chrome/browser automation.

1.14k repo starsObserved in 1 repos
Testing & Quality

tech-debt

Technical debt management - scan codebase for bad smells and create tracking issues

1.14k repo starsObserved in 1 repos
Testing & Quality

bug-analysis

Analyze software bugs for the PigeonPod project with a bugfix-first workflow. Use when users report broken behavior, regressions, incorrect results, crashes, data inconsistencies, sync/download failures, or ask for root-cause analysis, fix strategy, repro analysis, severity assessment, or regression-risk evaluation. Read current repository docs and code first, then use MCP tools including Context7 only when framework, library, API, or external-service behavior must be verified.

1.13k repo starsObserved in 1 repos
Testing & Quality

liveagent-code-review

Review an open GitHub pull request or the current local branch and working tree with parallel, independent reviewers and evidence-based validation. Use when the user asks for code review, invokes the Code Review action from Git Review, or explicitly mentions this skill.

1.13k repo starsObserved in 1 repos
Testing & Quality

doc-reviewer

Reviews recent code changes and checks if documentation needs updates. Reads all MD files in root and docs/ to identify stale or missing documentation. Use when completing features, before pushing, or when asked to "check docs", "review documentation", or "doc-reviewer".

1.13k repo starsObserved in 2 repos
Testing & Quality