
示例工程【免费下载链接】coursesAnthropics educational courses项目地址https://gitcode.com/GitHub_Trending/cours/courses点击查看免费下载本教程来自courses仓库中 Prompt evaluations 系列课程的第 9 课讲解如何在 promptfoo 中编写自定义模型评分器custom model-graded evaluator让一个 LLMClaude 3.5 Sonnet为另一个 LLM 生成的摘要打分。读完本文你将掌握promptfooconfig.yaml的完整配置、get_assert约定入口、带评分细则与少量示例的评分提示词设计以及通过promptfoo eval/promptfoo view运行和解读评测结果的全流程。1. 评测目标把复杂维基百科文章变成小学生能读懂的摘要本课的实战场景非常具体给定一篇篇幅长、术语密集的维基百科文章例如卷积神经网络的完整词条我们希望得到一个面向小学阶段、无技术背景读者的简短摘要例如Convolutional neural networks, or CNNs, are a special type of computer program that can learn to recognize images and patterns. They work a bit like the human brain, using layers of artificial neurons to process information. … Scientists and engineers use CNNs for all sorts of cool applications, like helping self-driving cars see the road, finding new medicines, or even teaching computers to play games like chess and Go.摘要好不好是典型的主观质量问题难以用简单的字符串匹配或代码规则衡量。因此本课的做法是用模型来给模型打分——定义一个自定义断言assertion让 Claude 依据一套明确的多维度评分细则rubric对摘要打分再把三个维度的分数平均作为该测试用例的最终得分。评分维度共三个每个维度 15 分Conciseness简洁性摘要是否尽可能精炼Accuracy准确性摘要是否与原文完全一致、无错误或误导Tone语气摘要是否适合无技术背景的小学读者三个维度取平均后以4.5/5作为通过阈值。整个实验的完整代码与数据都位于仓库目录 prompt_evaluations/09_custom_model_graded_prompt_foo/逐行讲解在 lesson.ipynb 中。它承接前两课的基础第 7 课的自定义评分器语法07_prompt_foo_custom_graders和第 8 课的模型评分概念08_prompt_foo_model_graded。2. 输入数据articles 目录中的 8 篇真实维基百科文本评测的第一步是准备输入数据。本课在articles/目录下提供了 8 个.txt文件每个文件包含一篇维基百科文章的完整正文文本例如 article1.txt大型语言模型词条、article4.txt 等。这些文件从几 KB 到几十 KB 不等内容长且技术性强正好用于检验摘要提示词的降复杂度能力。课程明确提醒8 个测试用例对真实评测而言远远不够。整套课程反复强调实际评测数据集至少应包含 100 条左右的样本才能对提示词质量形成有统计意义的判断。这里的 8 条仅用于教学演示。由于文章文本太长不可能把内容直接内联写进 YAML 配置文件promptfoo 提供了从文件加载变量的file://语法下一节详解。3. 被评测的三个提示词prompts.pyprompts.py 定义了三个待评测的提示词生成函数它们接受article变量并返回完整提示词def basic_summarize(article): return fSummarize this article {article} def better_summarize(article): return f Summarize this article for a grade-school audience: {article} def best_summarize(article): return f You are tasked with summarizing long wikipedia articles for a grade-school audience. Write a short summary, keeping it as concise as possible. The summary is intended for a non-technical, grade-school audience. This is the article: {article}三个提示词的能力梯度非常清晰basic_summarize只有一个动词指令better_summarize增加了面向小学读者的受众约束best_summarize则补充了任务定位为小学生总结长维基百科文章、输出要求尽量简洁和受众说明。课程特别说明这些提示词都是刻意写得很平庸的——没有遵循最佳实践例如没有加入完整示例目的是控制整个评测集运行时消耗的 token 数量。评测的价值就在于用数据证明即使不做大规模提示词工程仅仅是更明确的指令也能带来可测量的质量差异。4. 评测配置promptfooconfig.yaml 逐段解析promptfooconfig.yaml 是本次评测的完整配置description: Summarization Evaluation prompts: - prompts.py:basic_summarize - prompts.py:better_summarize - prompts.py:best_summarize providers: - id: anthropic:messages:claude-3-5-sonnet-20240620 label: 3.5 Sonnet tests: - vars: article: file://articles/article1.txt - vars: article: file://articles/article2.txt - vars: article: file://articles/article3.txt - vars: article: file://articles/article4.txt - vars: article: file://articles/article5.txt - vars: article: file://articles/article6.txt - vars: article: file://articles/article7.txt - vars: article: file://articles/article8.txt defaultTest: assert: - type: python value: file://custom_llm_eval.py逐字段拆解description评测的名称会显示在结果界面中便于区分多个评测项目。prompts以文件:函数名的语法引用 prompts.py 中定义的三个函数。promptfoo 会对每个提示词 × 每个测试用例做笛卡尔积式运行即 3 个提示词 × 8 篇文章 24 次生成调用。providers指定被评测的模型。这里使用anthropic:messages:claude-3-5-sonnet-20240620并通过label: 3.5 Sonnet起一个易读的显示名。anthropic:messages:前缀对应 Anthropic Messages API。tests定义 8 个测试用例每个用例只设置一个article变量。新语法点是file://articles/article1.txt——告诉 promptfoo 将变量的值设为该文件的文本内容从而避免把数万字符的维基百科正文内联进 YAML。defaultTest.assert对每一个测试用例统一生效的断言。type: python表示运行自定义 Python 断言value: file://custom_llm_eval.py指向断言实现文件。语法与第 7 课自定义评分器一致唯一的区别是这次的断言函数内部会再调用一次 Claude 来完成评分——即用模型评模型。5. 核心实现自定义模型评分器 custom_llm_eval.pycustom_llm_eval.py 是本课的灵魂文件包含两个函数get_assert与llm_eval。5.1get_assert()promptfoo 的约定入口promptfoo 在type: python断言中会自动查找名为get_assert的函数并传入两个参数output被评测模型这里是 3.5 Sonnet生成的实际输出即摘要文本context包含本次测试的变量context[vars]与生成该输出的提示词等信息的字典。promptfoo 允许该函数返回三种结果之一布尔值pass/fail、浮点数score或一个 GradingResult 字典。本课采用字典形式它必须包含三个属性pass布尔、score浮点、reason字符串解释。def get_assert(output: str, context, threshold4.5): article context[vars][article] score, evaluation llm_eval(output, article ) return { pass: score threshold, score: score, reason: evaluation }带注释的等价版本清晰展示了数据流def get_assert(output: str, context, threshold4.5): # 从 context 中取出当前测试用例对应的文章 article context[vars][article] # 把模型输出与文章原文交给 llm_eval得到分数与评估说明 score, evaluation llm_eval(output, article) # 返回 pass/score/reason 三元组供 promptfoo 判定与展示 return { pass: score threshold, score: score, reason: evaluation }注意threshold4.5是默认参数直接内联在get_assert中如需调阈值改这一处即可。5.2llm_eval()用 Claude 执行评分的完整逻辑llm_eval是真正干活的函数流程可概括为四步构造一份非常详尽的评分细则提示词rubric说明三个维度各自的打分标准通过 Anthropic Python SDK 调用 Claude 3.5 Sonnet 执行评分解析返回的 JSON计算三个数值维度的平均分返回平均分与模型的完整文本响应。完整实现如下评分提示词是核心务必逐字理解import anthropic import os import json def llm_eval(summary, article): Evaluate summary using an LLM (Claude). Args: summary (str): The summary to evaluate. article (str): The original text that was summarized. Returns: bool: True if the average score is above the threshold, False otherwise. client anthropic.Anthropic(api_keyos.getenv(ANTHROPIC_API_KEY)) prompt fEvaluate the following summary based on these criteria: 1. Conciseness (1-5) - is the summary as concise as possible? - Conciseness of 1: The summary is unnecessarily long, including excessive details, repetitions, or irrelevant information. It fails to distill the key points effectively. - Conciseness of 3: The summary captures most key points but could be more focused. It may include some unnecessary details or slightly overexplain certain concepts. - Conciseness of 5: The summary effectively condenses the main ideas into a brief, focused text. It includes all essential information without any superfluous details or explanations. 2. Accuracy (1-5) - is the summary completely accurate based on the initial article? - Accuracy of 1: The summary contains significant errors, misrepresentations, or omissions that fundamentally alter the meaning or key points of the original article. - Accuracy of 3: The summary captures some key points correctly but may have minor inaccuracies or omissions. The overall message is generally correct, but some details may be wrong. - Accuracy of 5: The summary faithfully represents the main gist of the original article without any errors or misinterpretations. All included information is correct and aligns with the source material. 4. Tone (1-5) - is the summary appropriate for a grade school student with no technical training? - Tone of 1: The summary uses language or concepts that are too complex, technical, or mature for a grade school audience. It may contain jargon, advanced terminology, or themes that are not suitable for young readers. - Tone of 2: The summary mostly uses language suitable for grade school students but occasionally includes terms or concepts that may be challenging. Some explanations might be needed for full comprehension. - Tone of 3: The summary consistently uses simple, clear language that is easily understandable by grade school students. It explains complex ideas in a way that is accessible and engaging for young readers. 5. Explanation - a general description of the way the summary is evaluated examples example This summary: summary Artificial neural networks are computer systems inspired by how the human brain works. They are made up of interconnected neurons that process information. These networks can learn to do tasks by looking at lots of examples, similar to how humans learn. Some key things about neural networks: - They can recognize patterns and make predictions - They improve with more data and practice - Theyre used for things like identifying objects in images, translating languages, and playing games Neural networks are a powerful tool in artificial intelligence and are behind many of the smart technologies we use today. While they can do amazing things, they still arent as complex or capable as the human brain. summary Should receive a 5 for tone, a 5 for accuracy, and a 5 for conciseness /example example This summary: summary Here is a summary of the key points from the article on artificial neural networks (ANNs): 1. ANNs are computational models inspired by biological neural networks in animal brains. They consist of interconnected artificial neurons that process and transmit signals. 2. Basic structure: - Input layer receives data - Hidden layers process information - Output layer produces results - Neurons are connected by weighted edges 3. Learning process: - ANNs learn by adjusting connection weights - Use techniques like backpropagation to minimize errors - Can perform supervised, unsupervised, and reinforcement learning 4. Key developments: - Convolutional neural networks (CNNs) for image processing - Recurrent neural networks (RNNs) for sequential data - Deep learning with many hidden layers 5. Applications: - Pattern recognition, classification, regression - Computer vision, speech recognition, natural language processing - Game playing, robotics, financial modeling 6. Advantages: - Can model complex non-linear relationships - Ability to learn and generalize from data - Adaptable to many different types of problems 7. Challenges: - Require large amounts of training data - Can be computationally intensive - Black box nature can make interpretability difficult 8. Recent advances: - Improved hardware (GPUs) enabling deeper networks - New architectures like transformers for language tasks - Progress in areas like generative AI The article provides a comprehensive overview of ANN concepts, history, types, applications, and ongoing research areas in this field of artificial intelligence and machine learning. /summary Should receive a 1 for tone, a 5 for accuracy, and a 3 for conciseness /example /examples Provide a score for each criterion in JSON format. Here is the format you should follow always: json {{ conciseness: number, accuracy: number, tone: number, explanation: string, }} /json Original Text: original_article{article}/original_article Summary to Evaluate: summary{summary}/summary response client.messages.create( modelclaude-3-5-sonnet-20240620, max_tokens1000, temperature0, messages[ { role: user, content: prompt }, { role: assistant, content: json } ], stop_sequences[/json] ) evaluation json.loads(response.content[0].text) # Filter out non-numeric values and calculate the average numeric_values [value for key, value in evaluation.items() if isinstance(value, (int, float))] avg_score sum(numeric_values) / len(numeric_values) # Return the average score and the overall model response return avg_score, response.content[0].text def get_assert(output: str, context, threshold4.5): article context[vars][article] score, evaluation llm_eval(output, article ) return { pass: score threshold, score: score, reason: evaluation }这份评分提示词中有几个值得反复琢磨的工程细节锚点式评分细则rubric anchors每个维度只给出 1/3/5 分的具体描述把什么是 1 分、什么是 5 分锚定成可操作的文本显著降低评分的随意性。少样本示例few-shot examplesexamples块给出了一个好摘要应得 5/5/5 分和一个专业但冗长的摘要应得 1/5/3 分作为参照告诉模型语气太技术化会扣分、过度详尽会扣简洁性分。示例让评分标准从抽象描述落地为具体形态。强制 JSON 输出要求模型严格按json{conciseness: number, accuracy: number, tone: number, explanation: string}/json的格式输出。代码中用{{和}}转义了 f-string 中的花括号。结构化输出的三重保险在messages中预先填入{role: assistant, content: json}相当于替模型起个头把输出直接引导到 JSON 标签内同时设置stop_sequences[/json]让模型在写完 JSON 后立即停止max_tokens1000给足输出空间。三者配合使json.loads(response.content[0].text)能稳定解析。temperature0评分任务追求确定性温度归零可尽量保证同一摘要多次评分的稳定性。平均分计算解析出的 JSON 字典中只有conciseness、accuracy、tone是数值explanation是字符串代码用isinstance(value, (int, float))过滤出数值字段再求平均天然排除了说明字段的干扰。返回完整评估文本llm_eval同时返回平均分和response.content[0].text即带 explanation 的完整评分输出后者被填入get_assert的reason字段——这样在 promptfoo 结果界面里可以看到模型给出该分数的完整理由便于人工审计。6. 运行评测与查看结果运行前需要先设置 Anthropic API 密钥README.md 中的说明export ANTHROPIC_API_KEYsk-ant-xxxx然后执行评测。README 中给出的最简命令是promptfoo evallesson.ipynb 中则使用npx直接调用最新版 promptfoo无需全局安装npx promptfoolatest eval运行完成后查看结果npx promptfoolatest view课程特别提醒这次评测会比较慢。原因在于每个测试用例都包含两轮模型调用——第一轮让 3.5 Sonnet 生成摘要3 个提示词 × 8 篇文章第二轮再让 3.5 Sonnet 为每份摘要评分同样 24 次。生成 评分的双重调用意味着 token 开销和耗时都是普通评测的两倍量级这也是前面提示词刻意保持简短的原因。7. 结果解读从终端到 Web 面板运行promptfoo eval后终端会输出 3 个提示词 × 8 篇文章的矩阵结果。随后promptfoo view会启动本地 Web 面板把每次运行的结果可视化点击矩阵中任意单元格的放大镜可以查看该次测试的详细评分信息。下图展示了一个典型的失败用例——某个摘要因为语气tone得分过低而未通过自定义评分函数Web 面板的顶行会汇总每个提示词的总体评分一目了然地比较三个提示词的表现面板顶部还提供得分分布图帮助可视化各提示词在全部输入上的表现差异课程对分布图的解读如下图中红色代表basic_summarize、蓝色代表better_summarize、绿色代表best_summarize。结果并不意外——best_summarize从未在任何输入上失败并且在所有输入上都高于另外两个提示词basic_summarize由于只给了总结这篇文章的模糊指令得分最低且频繁失败。这说明即便不引入复杂的提示词工程技巧仅仅是把任务目标、受众和输出要求写清楚就能在模型评分维度上产生可量化的显著提升。8. 设计要点与适用边界回顾整个案例这套自定义模型评分方案有五个值得借鉴的设计要点评分标准必须显式化把好摘要拆解为简洁性、准确性、语气三个可打分维度并为每个维度提供 1/3/5 分的锚点描述评分才可复现、可解释。用示例约束主观判断一对好/坏对照示例远比抽象描述更能稳定模型评分尺度是评分提示词性价比最高的一部分。结构化输出要双保险assistant角色预填充jsonstop_sequences[/json]保证解析稳定temperature0保证评分确定性。把解释带进结果reason字段携带完整评分文本使每次 pass/fail 都可追溯、可审计这是模型评分优于黑盒阈值的核心价值。注意成本与规模模型评分意味着生成与评分两轮调用token 开销翻倍同时 8 个用例的规模只够教学演示真实评测建议至少 100 条样本整套课程 prompt_evaluations/README.md 中反复强调这一点。需要说明的边界本课演示的评分对象是摘要但自定义模型评分器的模式完全通用——任何难以用代码规则衡量、需要主观判断的质量维度如文案语气、回答友好度、代码可读性、翻译自然度都可以套用同一套get_assert llm_eval结构。要复现本实验只需将 prompt_evaluations/09_custom_model_graded_prompt_foo/ 整个目录下载下来设置ANTHROPIC_API_KEY后依次运行npx promptfoolatest eval与npx promptfoolatest view即可。赞分享示例工程【免费下载链接】coursesAnthropics educational courses项目地址https://gitcode.com/GitHub_Trending/cours/courses点击查看免费下载相关推荐用 Promptfoo 评估 Claude 摘要质量claude-cookbooks Summarization Evaluation 套件实战指南用 Promptfoo 评估 Claude 摘要质量claude cookbooks Summarization Evaluation 套件实战指南 本指南以示例工程Claude 文档摘要实战指南从基础 Prompt 到摘要索引 RAG 与 Promptfoo 评估claude-cookbooksClaude 文档摘要实战指南从基础 Prompt 到摘要索引 RAG 与 Promptfoo 评估claude cookbooks 导读 本文基于 cl示例工程promptfoo 自定义 Python 评分器实战为 Claude 编写代码级评估Custom Graderspromptfoo 自定义 Python 评分器实战为 Claude 编写代码级评估Custom Graders 本教程来自 Anthropic 提示词评示例工程上一篇终极指南Atom编辑器的组件化设计与编程范式实践下一篇Lucky配置备份自动化定时上传至云存储灾难恢复方案创作声明:本文部分内容由AI辅助生成(AIGC),仅供参考