Testing & Quality skills

Browse reusable Agent Skills, each with a clear purpose and practical guidance.

eccv-experiments

Use when designing or auditing the experimental program of an ECCV paper — benchmark selection that survives a September conference, matched-substrate baseline fairness in the foundation-model era, ablations that isolate the claimed mechanism, qualitative failure evidence, and run sequencing toward a March freeze.

805 repo starsObserved in 1 repos
Testing & Quality

icassp-reproducibility

Use when strengthening ICASSP reproducibility across signal-processing modalities — pinning the scoring ruler for the paper's metric, dataset versions and splits, front-end/DSP settings, seeds, and compute, and mapping each claim to a checkable location, since ICASSP has no reviewed appendix and the four pages plus a public release must carry it.

805 repo starsObserved in 1 repos
Testing & Quality

icdm-experiments

Use when designing or auditing the empirical evaluation for an ICDM (IEEE International Conference on Data Mining) paper - mining-task definition, strong and fairly-tuned baselines, ablations that isolate the named mechanism, scalability curves that test scale claims, and discovery-validity checks that separate real findings from evaluation artifacts.

805 repo starsObserved in 1 repos
Testing & Quality

icdt-reproducibility

Use when strengthening the verifiability of an ICDT (International Conference on Database Theory) paper — the theory-venue analogue of reproducibility — covering complete and self-contained proofs, exact models and assumptions, claim-to-proof mapping, matching upper and lower bounds, consistency between the LIPIcs paper and the arXiv full version, and honest labeling of what is proved versus conjectured.

805 repo starsObserved in 1 repos
Testing & Quality

issta-experiments

Use when designing or auditing ISSTA experiments, covering real subject programs and benchmarks like Defects4J, fair tool-baseline configuration, bug-finding and coverage metrics, non-parametric comparison with effect sizes, equal-budget protocols, repeated runs, and matching evidence to the claim being made.

805 repo starsObserved in 1 repos
Testing & Quality

jmcb-robustness

Use when a Journal of Money, Credit and Banking (JMCB) result may be specification-, sample-, or inference-sensitive and you need to plan checks that each kill a specific threat. Builds a threat-mapped robustness suite; it does not re-run the core identification or write the prose.

805 repo starsObserved in 1 repos
Testing & Quality

mobicom-experiments

Use when designing or auditing the evaluation of a MobiCom submission — building real-device testbeds, choosing RF and channel measurement methodology, injecting realistic mobility and interference, profiling energy on hardware, picking tuned baselines, and reporting distributions so wireless reviewers see where the mechanism wins and breaks.

805 repo starsObserved in 1 repos
Testing & Quality

neurips-experiments

Use when stress-testing NeurIPS experimental evidence, including baselines, ablations, data splits, compute, negative results, real-world use, and claim-to-evidence calibration.

805 repo starsObserved in 1 repos
Testing & Quality

oopsla-reproducibility

Use when hardening an OOPSLA paper's empirical claims to the SIGPLAN Empirical Evaluation Guidelines — managed-runtime measurement discipline, warmup and variance reporting, corpus and benchmark provenance, environment pinning, and a Data-Availability Statement that the eventual artifact can actually honor.

805 repo starsObserved in 1 repos
Testing & Quality

rt-replication-package

Use before final submission or on acceptance to assemble and validate the Data-Editor replication package against the target venue's data-and-code policy — master script, pinned environment, README/roadmap, restricted-data plan, and a script-to-exhibit output map, with a checklist that catches the numbers a Data Editor would fail to reproduce. Reads the venue policy live from the pack's source-map.

805 repo starsObserved in 1 repos
Testing & Quality