Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

aliyun-wan-r2v-test

Minimal reference-to-video smoke test for Model Studio Wan R2V.

396 repo starsObserved in 1 repos
Testing & Quality

analyzing-shaft-failures

Use when analyzing SHAFT Allure results, Doctor reports, trace evidence, healer output, flaky locator/wait/assertion failures, retries, or test-fix recommendations.

395 repo starsObserved in 1 repos
Testing & Quality

choosing-shaft-locators

Use when creating, reviewing, refactoring, repairing, or generating SHAFT web/mobile locators, smart locators, ARIA locators, XPath/CSS replacements, or codegen element identifiers.

395 repo starsObserved in 1 repos
Testing & Quality

flaky-test-stabilizer

Diagnose and stabilize intermittent SHAFT tests, especially TestNG parallelism, shared state, files, drivers, timing, and external dependencies.

395 repo starsObserved in 1 repos
Testing & Quality

recording-shaft-tests-with-mcp

Use when recording browser, Playwright, mobile, Appium Inspector, or user-performed flows into SHAFT tests with MCP Capture, replay, code blocks, or codegen insertion.

395 repo starsObserved in 1 repos
Testing & Quality

regression-test

Add a regression test for an already-fixed PHPStan bug given a GitHub issue number

395 repo starsObserved in 1 repos
Testing & Quality

verifying-and-applying-shaft-changes

Use when reviewing, previewing, applying, guardrail-checking, or verifying generated SHAFT Java before or after inserting it into a repository, especially the coding-partner diff/apply/verify loop from IntelliJ.

395 repo starsObserved in 1 repos
Testing & Quality

writing-shaft-tests

Use when writing, reviewing, or repairing SHAFT Java tests, page objects, API tests, mobile tests, CLI/DB tests, assertions, waits, or TestNG/JUnit/Cucumber scenarios.

395 repo starsObserved in 1 repos
Testing & Quality

complexa-evaluate-pdbs

Standalone evaluation of an existing PDB directory with Proteina-Complexa. Use this skill whenever the user wants to "evaluate PDB files", "re-fold these designs", "compute interface pAE", "compute i_pLDDT for a folder", "run AF2 / RF3 / ESMFold on my designs", "score binder candidates", "designability of this folder", "scRMSD for designs", "motif RMSD for these PDBs", "complexa analysis", "complexa evaluate from a PDB directory", "evaluate from pdb dir", or score third-party outputs (BindCraft, AlphaProteo, RFdiffusion, hand-curated decoys). It picks the correct `evaluate_*.yaml` config, wires `++dataset.pdb_dir` and the folding backend, runs `complexa analysis` (the evaluate → analyze chain), parses the result CSV, reports pass-rates against the right `result_type` thresholds, and emits a replayable `eval_manifest.json`. Reach for this skill before hand-rolling refolding scripts.

393 repo starsObserved in 1 repos
Testing & Quality

story-review

多视角对抗式审查。full/lean 模式在已部署 reviewer agents 时并行 spawn;缺失/异常 agents 或 spawn 失败时自动降级 solo,参考文件不可读时使用内置 rubric fallback。 触发方式:/story-review、/审查、「审查一下」「帮我审一下」

393 repo starsObserved in 1 repos
Testing & Quality