Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

aivisspeech-sentry-triage

AivisSpeech エディタの Sentry issue を調査し、修正すべき Electron / Vue 側の不具合と、ローカル環境・ブラウザ実装・外部通信由来のノイズを切り分けるためのスキルです。Sentry 側で既知ノイズを整理する作業や、src/domain/sentryEventFilter.ts と関連テストを更新して既知ノイズを送信前に破棄する作業で使用します。

457 repo starsObserved in 1 repos
Testing & Quality

host_bash

设备上跑只读 shell 命令做诊断探索(沙箱化 / read-only policy)

457 repo starsObserved in 1 repos
Testing & Quality

run-evaluation

Run a VLA model evaluation against a simulation benchmark. Use this skill whenever the user wants to evaluate, benchmark, test, or run a model on a sim environment — even if they say it casually like 'try OpenVLA on LIBERO' or 'get me CALVIN scores'. Covers the full workflow: serving the model, launching the benchmark, sharding for speed, merging results, and interpreting output.

456 repo starsObserved in 1 repos
Testing & Quality

E2E Test Builder

Create Playwright E2E tests using Page Object Model pattern with database isolation

454 repo starsObserved in 2 repos
Testing & Quality

shopify-app-store-review

Run a pre-submission compliance check against your Shopify app's codebase. Reviews App Store requirements and surfaces likely issues before you submit for official review.

454 repo starsObserved in 1 repos
Testing & Quality

animation-review

Reviews Godot Animation implementation for known pitfalls. Triggers AFTER implementation, when code involves AnimationPlayer, AnimationTree, AnimatedSprite2D, SpriteFrames, AnimationNodeStateMachine, BlendSpace, OneShot, callback_mode_process, or animation playback control (play/travel/start). Do NOT use this skill for planning or teaching — only for post-implementation review.

453 repo starsObserved in 1 repos
Testing & Quality

gdtoolkit

Lint and format GDScript files using gdtoolkit (gdlint + gdformat). Use after writing or modifying .gd files, when asked to check code style, fix lint errors, format code, or set up linting configuration. Also use when gdlint/gdformat errors appear in output and need diagnosis. Does NOT require Godot — runs as a standalone Python tool.

453 repo starsObserved in 1 repos
Testing & Quality

gdunit-driver

Run gdUnit4 unit tests and parse results into structured output. Use this skill after writing or modifying code to verify correctness via unit tests, when diagnosing test failures, or when writing new test files. Triggers: "run tests", "test fails", "write a test", any gdUnit4/unit test mention. Supports both GDScript (.gd) and C# (.cs) test files.

453 repo starsObserved in 1 repos
Testing & Quality

gm-evaluate

Evaluate the current tag's quality: enforce the playable-closed-loop gate, maintain a single cross-tag e2e/ suite that always reflects the current game (add tests for new mechanics, prune tests for mechanics this tag deliberately removed), and reason about gameplay quality. Independent from the build process — fresh perspective on the final product. Explicit invocation only — use /gm-evaluate.

453 repo starsObserved in 1 repos
Testing & Quality

gm-fixgap

Fix gaps identified by the Evaluator. Generates GAP.md from evaluation.json, dispatches workers to address critical/major issues, then runs one final verify+review pass. Unlike gm-build (PLAN.md-driven), gm-fixgap is GAP.md-driven. Explicit invocation only — use /gm-fixgap.

453 repo starsObserved in 1 repos
Testing & Quality