Loading...
Loading...
Found 1,893 Skills
Automated reproduction of comprehensive model evaluation benchmarks following the Benchmark Suite V3. Auto-activates for model benchmarking, comparison evaluation, or performance testing between AI models.
INVOKE THIS SKILL for LLM-as-judge evaluation workflows on Arize: creating/updating evaluators, running evaluations on spans or experiments, tasks, trigger-run, column mapping, and continuous monitoring. Use when the user says: create an evaluator, LLM judge, hallucination/faithfulness/correctness/relevance, run eval, score my spans or experiment, ax tasks, trigger-run, trigger eval, column mapping, continuous monitoring, query filter for evals, evaluator version, or improve an evaluator prompt.
Valuation methodology framework covering absolute (DCF / DDM / SOTP) and relative (PE-Band / PB-ROE / EV-EBITDA / PS) approaches — when to use each, pros/cons, common pitfalls, and practical application with Longbridge data. Triggers: "估值方法", "估值方法论", "DCF", "DDM", "SOTP", "PE估值", "EV/EBITDA", "绝对估值", "相对估值", "估值框架", "估值方法論", "絕對估值", "相對估值", "valuation methodology", "DCF model", "DDM", "SOTP", "PE band", "EV EBITDA", "absolute valuation", "relative valuation", "valuation framework".
Script a Roblox experience in Luau: get services, create and parent Instances, connect events, run server Scripts vs client LocalScripts, and communicate across the client/server boundary with RemoteEvents/RemoteFunctions (server-authoritative). Use when building or debugging Roblox Studio scripts — when the user mentions Roblox, Luau, services, RemoteEvent, Instance.new, PlayerAdded, or client vs server. For saving player data use roblox-datastores.
Build evaluation frameworks for agent systems. Use when testing agent performance, validating context engineering choices, or measuring improvements over time.
Runs a 25-point quality control test on a supplier sample before mass production. Catches the defects that turn into 18% return rates and stranded inventory. Produces a back-message to the supplier with specific fixes required before greenlighting the PO. Use when a user mentions Alibaba sample, supplier sample, sample QC, or before mass production. Trigger phrases: "Amazon supplier sample", "Alibaba sample evaluation", "sample QC", "before mass production", "supplier quality test". Works with zero tools.
Evaluate and score film and television scripts from three dimensions: ideological value, artistic merit, and viewability. Suitable for quality assessment during script development, determining revision directions, and pre-project approval review.
This skill should be used when the user asks to "evaluate agent performance", "build test framework", "measure agent quality", "create evaluation rubrics", or mentions LLM-as-judge, multi-dimensional evaluation, agent testing, or quality gates for agent pipelines. Part of the context engineering skill suite — also activates when the user mentions "context engineering" or "context-engineering" in the context of measuring agent effectiveness.
Systematic framework for evaluating scholarly and research work based on the ScholarEval methodology. This skill should be used when assessing research papers, evaluating literature reviews, scoring research methodologies, analyzing scientific writing quality, or applying structured evaluation criteria to academic work. Provides comprehensive assessment across multiple dimensions including problem formulation, literature review, methodology, data collection, analysis, results interpretation, and scholarly writing quality.
Use this when you need to EVALUATE OR IMPROVE or OPTIMIZE an existing LLM agent's output quality - including improving tool selection accuracy, answer quality, reducing costs, or fixing issues where the agent gives wrong/incomplete responses. Evaluates agents systematically using MLflow evaluation with datasets, scorers, and tracing. Covers end-to-end evaluation workflow or individual components (tracing setup, dataset creation, scorer definition, evaluation execution).
Discounted cash flow valuation and intrinsic value analysis for public companies. Use when the brief asks for DCF, fair value, intrinsic value, price target, undervalued or overvalued analysis, or "what is this company worth?"
Amazon Bedrock AgentCore Evaluations for testing and monitoring AI agent quality. 13 built-in evaluators plus custom LLM-as-Judge patterns. Use when testing agents, monitoring production quality, setting up alerts, or validating agent behavior.