Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

llm-sast-scanner

General-purpose Static Application Security Testing (SAST) skill for code vulnerability analysis. Trigger when the user asks to: "analyze code for vulnerabilities", "review code security", "find security bugs", "do a SAST scan", "check for [vulnerability type] in code", "audit source code", or requests a security code review of any language or framework. Covers 34 vulnerability classes across web, API, auth, mobile, and logic layers.

273 repo starsObserved in 2 repos
Testing & Quality

renderdoc-mcp

Analyze RenderDoc GPU frame captures with renderdoc-mcp MCP tools. Use when Codex needs to inspect .rdc captures, diagnose black screens or visual artifacts, explain frame structure, inspect specific draw calls, or investigate GPU rendering and performance issues.

272 repo starsObserved in 1 repos
Testing & Quality

sentry-audio-error-review

Review unresolved Sentry audio-request issues from the last 24 hours and report any whose exception_type looks mis-categorized against hypertts_addon/errors.py

272 repo starsObserved in 1 repos
Testing & Quality

benchmark-model

Scaffold the setup + eval scripts to run one model on the TabArena benchmark cluster. Use this skill whenever a maintainer wants to benchmark an already-integrated model — e.g. "benchmark TabM", "run Nori on the cluster", "create a setup/eval script for DenseLight", "I want to launch <model> on TabArena and evaluate it". Generates a single `tmp_scripts/run_<model>.py` with `setup` and `eval` subcommands that share one benchmark_name + paths (so they cannot drift), auto-filling GPU/CPU, eval subsets, the install reminder, foundation-model prefetch from the model's registry `info.py`, and — mandatory for GPU models — the VRAM budget (`fake_memory_for_estimates`; asks the user when the partition's VRAM is not inferable). Complements `add-model` (integrate a model) and `upload-method` (publish its results).

271 repo starsObserved in 1 repos
Testing & Quality

local-diff-review

基于本地 git diff 的代码审查技能。用于在提交 PR 前先做自检评审,严格复用 pr-review 的三维度标准(代码质量标准、FastGPT 风格规范、常见问题清单),支持审查未暂存改动、已暂存改动、某次提交、或当前分支相对基线分支的全部改动。

271 repo starsObserved in 1 repos
Testing & Quality

motel-debug

Debug applications with motel, a local OpenTelemetry ingest and query server. Use when the user wants runtime-evidence debugging with traces or logs, wants temporary debug instrumentation that can be removed later, or needs a repo wired to send OTLP/HTTP telemetry to a local motel server. If the target repo uses Effect or @effect/*, also read references/effect.md.

270 repo starsObserved in 1 repos
Testing & Quality

deep-analyse

Thorough code analysis without polluting main context

269 repo starsObserved in 1 repos
Testing & Quality

i18n-validator

Validates multi-language consistency in the TRANSLATIONS dictionary of a plugin. Use to check if any language keys are missing or if translations need updating.

269 repo starsObserved in 1 repos
Testing & Quality

magic-spell

When the user's message contains ::__ASTRO_CLAW_SPELL__::, output a fixed status block. Used to verify skill loading works.

269 repo starsObserved in 1 repos
Testing & Quality

playwright-openwebui

Use when inspecting UI bugs in OpenWebUI plugins, taking screenshots of plugin output, capturing console errors, testing Action/Filter/Pipe plugins in the chat interface, or verifying plugin installation in the Admin panel. Triggered by: plugin UI bug, Action HTML output, screenshot, console error, plugin not working, login, admin panel, chat test, function install.

269 repo starsObserved in 1 repos
Testing & Quality