validity-probing
Challenge construct validity — does benchmark measure claimed capability? — 3 benchmarks, 40 papers, 30 web searches
Browse reusable Agent Skills, each with a clear purpose and practical guidance.
Challenge construct validity — does benchmark measure claimed capability? — 3 benchmarks, 40 papers, 30 web searches
Enumerate 3-5 values per parameter including extremes
Sobol variance decomposition — compute first-order and total-order sensitivity indices to quantify each parameter's contribution to output variance.
Identify orthogonal axes along which a method's validity might vary. Ensures axes are independent, measurable, and span the relevant parameter space.
Map 15 vital relations between concepts
Simplified crystallization strategy for users who have a general research direction (e.g., "I'm interested in LLM reasoning") but lack specificity. Simplifies actor profiling and landscape reconnaissance, then proceeds through direction narrowing, obstacle analysis, goal decomposition, and north-star synthesis. Use when the user's first message reveals a general area but not a specific problem.
Compute criteria weights using a specified elicitation method (AHP, Swing, BWM, MACBETH, or Simos).
Identify patent coverage gaps — feature cross-matrix blank areas revealing unprotected technology combinations. Budget: 150 patent families, 10 claim parses, 60 web searches.
Identify matrix regions not covered by existing methods
Apply Rittel's 10 criteria to determine if the problem is tame, complex, or wicked, and adjust research strategy accordingly.