This skill should be used for advanced LLM evaluation: LLM-as-judge systems, direct scoring, pairwise comparison, rubric calibration, evaluator bias mitigation, confidence scoring, and automated quality assessment.
Third-party extension
This extension is indexed and displayed by LangBot; copyright belongs to its original author and it is not officially maintained on LangBot Space. Before installing or using it, review its open-source license, README and security risks.
Loading...
Use Advanced Evaluation in Feishu/Lark bots, DingTalk bots, WeCom bots, WeChat bots, QQ bots, Slack bots, Discord bots, Telegram bots, LINE Bots, and Matrix Bots.
Host your bot on LangBot Cloud and connect this extension to the collaboration platforms your team already uses.