Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

icml-experiments

Use when stress-testing ICML experimental evidence before submission or rebuttal, including strong tuned baselines, mechanism-isolating ablations, seed variance and confidence intervals, compute disclosure, data leakage and split construction, reproducibility, negative results, and fit to ICML soundness, originality, and significance scoring.

805 repo starsObserved in 2 repos
Testing & Quality

icsme-reproducibility

Use when strengthening IEEE ICSME reproducibility and open-science evidence, covering the data-availability statement, anonymized-but-runnable artifacts, mining and LLM provenance pinning, claim-to-evidence mapping, honest degrees of reproducibility, and consistency between the paper and the artifact ahead of the ROSE-Festival IEEE badges.

805 repo starsObserved in 2 repos
Testing & Quality

ijcai-experiments

Use when designing or auditing IJCAI or IJCAI-ECAI experiments, baselines, ablations, statistical evidence, hyperparameter reporting, compute descriptions, dataset handling, ethics risks, and reproducibility evidence for AI papers.

805 repo starsObserved in 2 repos
Testing & Quality

iros-artifact-evaluation

Use when packaging IROS evidence for a skeptical reviewer even though IROS runs no formal artifact track — the robotics artifact stack of hardware ledger, logs, configs, code, and data, what an embodied-systems reviewer opens first, review-time anonymity versus acceptance-time public release, and making a claim auditable not asserted.

805 repo starsObserved in 2 repos
Testing & Quality

isca-experiments

Use when designing or auditing the evaluation of an ISCA paper — pinning simulator fidelity to the claims it must carry, documenting gem5-class configurations and sampling choices, selecting workload suites that represent the claim's domain, tuning baselines in good faith, and separating architectural effect from modeling artifact.

805 repo starsObserved in 2 repos
Testing & Quality

issta-reproducibility

Use when strengthening ISSTA reproducibility and verifiability evidence, covering pinned subject programs and benchmark versions, random seeds and timeout budgets, non-determinism disclosure for fuzzing and analysis, claim-to-evidence traceability, tool availability statements, and keeping the artifact consistent with the paper's tables.

805 repo starsObserved in 2 repos
Testing & Quality

jcadcg-reproducibility

在为《计算机辅助设计与图形学学报》(Journal of Computer-Aided Design & Computer Graphics, JCAD&CG) 提升论文可复现性与实验可重复时调用。覆盖图形学/CAD 特有的可复现要点:固定随机种子、记录 GPU/驱动/渲染器版本、锁定网格/点云/纹理/视频数据集、固定渲染与几何管线参数(采样数、光照、相机、分辨率)、复现包冒烟测试、几何误差与渲染质量指标的计算脚本对齐、以及大数据的稳定托管。适用于让外审与读者能重跑出论文表图、避免"结果无法复现"质疑的场景。

805 repo starsObserved in 2 repos
Testing & Quality

jf-robustness

Use when planning or auditing the robustness, sensitivity, and multiple-testing battery for a The Journal of Finance (JF) manuscript. Decides which checks are load-bearing; it does not design the main test.

805 repo starsObserved in 2 repos
Testing & Quality

jimf-robustness

Use when a Journal of International Money and Finance (JIMF) result must be shown stable across samples, regimes, measures, and inference choices. Designs the robustness layer mapped to international-finance threats; it does not establish the design (jimf-identification / jimf-empirical-design).

805 repo starsObserved in 2 repos
Testing & Quality

micro-experiments

Use when designing or auditing the evaluation of a MICRO paper — choosing the right instrument on the ladder from analytical model to cycle-level simulator to RTL to silicon, tuning baselines the PC will respect, selecting workload suites, running ablations and sensitivity sweeps, and reporting geomeans with full overhead accounting.

805 repo starsObserved in 2 repos
Testing & Quality