Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

sigcomm-artifact-evaluation

Use when packaging an ACM SIGCOMM paper's code, traces, topologies, and configuration for the artifact-evaluation committee — choosing ACM badges (Artifacts Available, Evaluated, Results Reproduced) as claim calibration, building downscaled topologies and trace substitutes, and making a networking testbed result rebuildable by a reviewer.

805 repo starsObserved in 1 repos
Testing & Quality

sigcomm-experiments

Use when designing or auditing ACM SIGCOMM experiments — climbing the evidence ladder from microbenchmarks to testbed and trace replay to deployment, choosing baselines that match the deployed state of the art, reporting tail percentiles and variance over repeated trials, mapping break points, and holding the fair variable fixed.

805 repo starsObserved in 1 repos
Testing & Quality

sigmetrics-experiments

Use when designing or auditing ACM SIGMETRICS evaluations, covering theorem-plus-validation rigor, stating and testing modeling assumptions, analysis-vs-simulation agreement, real workloads and traces, fairly tuned baselines, statistics and confidence intervals for stochastic systems, learning guarantees, measurement provenance, and matching evidence to the shape of each performance claim.

805 repo starsObserved in 1 repos
Testing & Quality

dev-proxy

Simulate API failures, mock responses, test rate limiting, and analyze API traffic using Dev Proxy's plugin-based proxy engine. WHEN: 'mock API responses', 'simulate API errors', 'test rate limiting', 'test error handling', 'mock OpenAI responses', 'test AI app', 'analyze API usage', 'configure Dev Proxy', 'install Dev Proxy', 'set up Dev Proxy', 'use Dev Proxy in CI/CD', 'chaos testing for APIs', 'run Dev Proxy in background', 'detached mode', 'run Dev Proxy detached'.

801 repo starsObserved in 1 repos
Testing & Quality

search-benchmark

Generate a search quality benchmark for the AI Registry. Generates ground truth from the registry's assets, runs 100+ queries against the semantic search API, evaluates results using NDCG@10/MRR/Recall, and produces a markdown report. Use when you want to measure search quality after changes to the scoring algorithm, embedding model, or indexed content.

800 repo starsObserved in 1 repos
Testing & Quality

nemo-mbridge-perf-cuda-graphs

Validate and use CUDA graph capture in Megatron Bridge, including local full-iteration graphs and Transformer Engine scoped graphs for attention, MLP, and MoE modules.

798 repo starsObserved in 6 repos
Testing & Quality

nemo-mbridge-perf-memory-tuning

Techniques for reducing peak GPU memory in Megatron Bridge — expandable segments, PEFT + SP input re-gather, parallelism resizing, activation recompute, CPU offloading constraints, and common OOM fixes.

798 repo starsObserved in 6 repos
Testing & Quality

nemo-mbridge-perf-expert-parallel-overlap

Validate and use MoE expert-parallel communication overlap in Megatron-Bridge, including overlap_moe_expert_parallel_comm, delay_wgrad_compute, and flex dispatcher backends such as DeepEP and HybridEP.

798 repo starsObserved in 4 repos
Testing & Quality

aiq-release-qa

Use when validating an AI-Q change before opening or merging a PR — choosing and running the right Python, frontend, docs, or eval checks for the surfaces you touched instead of one fixed command list.

795 repo starsObserved in 1 repos
Testing & Quality

debug-formatter

Debug a SQL formatter bug. Use when the user reports incorrect formatting output — wrong whitespace, misplaced comments, blank lines, etc.

793 repo starsObserved in 1 repos
Testing & Quality