Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

anmath-referee-strategy

Use when anticipating how an expert Annals of Mathematics referee will verify a proof — stress-testing the argument for gaps, checking external dependencies, and pre-empting the questions a careful reader will ask. Hardens the manuscript before submission; does not write the rebuttal (anmath-revision).

805 repo starsObserved in 2 repos
Testing & Quality

ase-experiments

Use when designing or auditing the evaluation of an ASE (IEEE/ACM Automated Software Engineering) paper, covering real subject systems, fair runnable tool baselines, task-matched effectiveness metrics, ablations that isolate a learned component, oracle and correctness validation, contamination-aware LLM handling, and provenance for mining.

805 repo starsObserved in 2 repos
Testing & Quality

asplos-experiments

Use when designing or auditing the evaluation of an ASPLOS paper — choosing among real silicon, FPGA prototypes, and simulators with cycle-accuracy caveats stated, selecting workload suites and baselines that hold up across three communities, attributing wins via ablation, and reporting energy, area, and overhead honestly.

805 repo starsObserved in 2 repos
Testing & Quality

atc-experiments

Use when designing or auditing the evaluation of an ATC (ACM SIGOPS Annual Technical Conference, formerly USENIX ATC) systems paper — matching evidence to the claim with real testbeds, fair baselines, end-to-end plus microbenchmark results, tail-latency and variance reporting, workload realism, and honest cost accounting.

805 repo starsObserved in 2 repos
Testing & Quality

cav-artifact-evaluation

Use when packaging a CAV (Computer Aided Verification) artifact for the Artifact Evaluation Committee (AEC), covering the three badges (Available / Functional / Reusable), the smoke-test and full-review phases, ≥2 AEC reviewers per artifact, DOI-issuing archives, verification-tool packaging (solvers, benchmarks, seeds, resource limits, proof witnesses), and the fact that AE is invited, post-notification, and non-conditional.

805 repo starsObserved in 2 repos
Testing & Quality

cav-experiments

Use when designing or auditing a CAV (Computer Aided Verification) empirical evaluation, covering standard benchmark sets (SV-COMP/SMT-COMP/HWMCC/VNN-COMP), fair baseline solvers with pinned versions and equal resource limits, timeout-dominated comparisons, soundness cross-checks and proof witnesses, cactus/scatter reporting, and matching evidence to the shape of each verification claim.

805 repo starsObserved in 2 repos
Testing & Quality

ccs-reproducibility

Use when strengthening ACM CCS reproducibility evidence, including the artifact-availability posture, threat-model-to-evidence mapping, attack reproduction steps, defense overhead measurement, measurement-dataset provenance, environment and version pinning, and honest justification when artifacts cannot be shared.

805 repo starsObserved in 2 repos
Testing & Quality

cjc-reproducibility

在为投向《计算机学报》(Chinese Journal of Computers, CJC) 的实证或系统类长文构建可复现性与实验可重复保障时调用。覆盖实验环境与依赖固定、随机种子与统计稳定性、数据来源与预处理的可追溯记录、基线与超参数的公平呈现、机器学习/数据挖掘研究的数据泄漏与污染防范、复现脚本与论文表图的一一对应,以及在双盲外审下如何组织可核验又不暴露作者身份的复现材料。适用于让计算机全学科中文原创研究经受三审专家对结果可信度的审查。

805 repo starsObserved in 2 repos
Testing & Quality

colt-artifact-evaluation

Use when deciding what evidence package a COLT (Conference on Learning Theory) paper needs, given that COLT runs no artifact-evaluation track or badges — the proof appendix is the artifact. Covers proof-verification passes, optional code companions for numerics, formalization aids, and post-acceptance release of scripts.

805 repo starsObserved in 2 repos
Testing & Quality

csj-reproducibility

在为《计算机科学》(Computer Science, JSJKX) 提升实验可复现性与可重复性时调用。本刊是计算机全学科中文综合月刊(CCF 会刊、B 类、T2 级),单盲审稿,未见独立 artifact 徽章制度(待核实),故复现工作主要服务于让外审专家更快确认方法可信。技能覆盖环境/依赖/随机种子/数据划分固定、数据泄漏与污染防范、复现包结构、外链数据与校验和、可复现性自查清单。适用于让稿件的主结果可被他人从原始数据一键复现、经得起外审追问的场景。

805 repo starsObserved in 2 repos
Testing & Quality