Testing & Quality skills
Browse reusable Agent Skills, each with a clear purpose and practical guidance.
gini-bug-report
File a locally-captured, already-redacted Gini crash report as a GitHub issue, with the user's consent. Reads the pending crash queue and delegates the actual filing to the github-issues skill.
e2e-conventions
When to write e2e tests, where to put them, and how to verify them. Apply to any task touching UI, filters, forms, or interactions.
qtpass-fixing
Bug fixing workflow for QtPass - find, fix, test, PR
qtpass-localization-audit
QtPass localization audit - structural checks on .ts files (placeholders, HTML balance, mnemonics, mixed-script artifacts)
qtpass-testing
Comprehensive guide for QtPass unit testing with Qt Test
spec-review
Independent, multi-model review of a design or spec **before** any code is written, for the microsoft/winappcli repo. Activate when a contributor asks to "review this spec", "review my design", "review this design doc", "validate this approach", "should we build this", "spec review", "design review", or "feature review". Fans out parallel sub-agents — each doing its OWN research against the real codebase and ecosystem rather than trusting the spec — covering necessity & scope, approach & alternatives, feasibility vs reality, risks & unknowns, DX & user impact, and a different-model-family cross-check. Emits a decision-oriented recommendation (proceed / proceed-with-changes / reconsider) to stdout. This is the PRE-CODE companion to the pr-review skill (which reviews code already written); use spec-review at the design/spec stage, not on an implemented diff. Does NOT write code or edit the spec.
winapp-troubleshoot
Diagnose and fix common Windows app packaging, signing, identity, and SDK errors. Use when encountering errors with MSIX packaging, certificate signing, Windows SDK setup, or app installation.
challenge-troubleshoot
Use when something in the Simulation Challenge pipeline is misbehaving — auth errors, agent disconnects, jobs stuck in Pending, jobs ending in Failed, drain frames. Maps a symptom to its likely cause and the next command to run.
check-inference
Probe a model inference WebSocket server (e.g. `serve_policy`) and validate the response — using the `geniesim benchmark check-inference` CLI verb, which wraps the benchmark package's `check_inference.py`. Trigger: When the user asks to "check inference", "校验模型推理", "test inference server", "verify policy server", "ping the model", or provides an IP/port and wants to confirm a serve_policy / WebSocket inference server is working before running benchmarks.