A CodeSignal learning project that teaches how to evaluate LLM prompts — not only how to write them.
Learners run the same prompt template against an input across multiple independent LLM calls, then score outputs against an optional expected answer using simple metrics.
- Done — Prompt template + input + 1–5 independent runs → collect outputs
- Done — Optional expected answer + metrics + mean/min/max
- Done — Prompt A vs Prompt B under shared conditions (winner by mean)
- Done — Evaluation across multiple test cases (overall + per-case scores)
- Next — Charts / distributions, or model/provider comparison
git clone --recurse-submodules <this-repo-url>
cd learn_cosmo-prompteval
npm install
cp .env.example .env
cp session.config.example.json session.config.jsonFill in .env with the API key (and optional *_BASE_URL) for the provider you want. Choose the model in session.config.json as provider/model-id:
anthropic/claude-sonnet-4-6— needsANTHROPIC_API_KEY, optionalANTHROPIC_BASE_URLopenai/gpt-5.6-luna— needsOPENAI_API_KEY, optionalOPENAI_BASE_URLgoogle/gemini-3.6-flash— needsGOOGLE_API_KEY, optionalGOOGLE_BASE_URL(gemini/…also routes to Gemini)~deepseek/deepseek-v4-flash-latest— needsDEEPSEEK_API_KEYandDEEPSEEK_BASE_URL(deepseek/…anddeepseek-ai/…also route here; uses the OpenAI SDK). If both DeepSeek vars are unset, it reusesOPENAI_API_KEY/OPENAI_BASE_URL(production proxy hack).
session.config.json is separate from .env. It is local (not checked in) and holds session defaults, not secrets:
model(optional) —provider/model-id(defaultanthropic/claude-sonnet-4-6); must be listed inallowedModelsallowedModels(optional) — picker list ofprovider/model-idrefs (defaults to Anthropic, OpenAI, Gemini, and DeepSeek examples above)allowUserModelSelection(optional) — whentrue, show a model picker and let the saved eval session overridemodelwith an entry fromallowedModels(defaultfalse)allowCompare(optional) — whentrue, show “Compare with another prompt” so learners can A/B two prompts (defaultfalse)allowedMetricIds(optional) — metrics shown in the picker and accepted by the API. When omitted, the original Course 1 metrics remain unchanged:exact-match,exact-match-ci,contains,string-similarity, andword-overlap-f1. Opt in to newer validation withregex-match,valid-json, and/orllm-judge.llmJudgeModel(optional) — fixedprovider/model-idused only byllm-judge. The server controls this value and the UI displays it to learners. When omitted, the generation model is reused for backward compatibility.features.promptTemplating(optional) — enables reusable named blanks, shared examples, and a filled-in prompt preview. Missing configuration keeps the current Course 1 UI and{{input}}behavior unchanged.maxConcurrency(optional) — max in-flight LLM calls during an evaluation (default4, range 1–50). Set to1for serial.defaults(optional) —runssets the initial run count whileminRuns,maxRuns,minCases, andmaxCasesset the editable limits (each 1–5)initialSession(optional) —promptA,promptB, andcases(input/expectedAnswer)
Without session.config.json, prompts and cases start empty and the UI uses the built-in 1–5 limits. Copy session.config.example.json to prefill the capital-city demo.
For example, a later course can enable every validation type without changing Course 1:
{
"llmJudgeModel": "anthropic/claude-sonnet-4-6",
"allowedMetricIds": [
"exact-match",
"regex-match",
"valid-json",
"llm-judge"
]
}regex-match treats Expected Answer as a regular expression. valid-json is a deterministic function checker and does not require an expected answer. llm-judge makes a second call to llmJudgeModel for each generated output and requires an expected answer. Function checkers are registered in code and enabled by ID; configuration never executes arbitrary JavaScript.
Later-course tasks can let the prompt template define the fields shown in every case without changing Course 1:
{
"allowCompare": false,
"features": {
"promptTemplating": {
"enabled": true,
"templateEditable": true,
"showPreview": true,
"dynamicFields": true,
"fields": [
{ "name": "context", "label": "Context" },
{ "name": "input", "label": "Input" },
{ "name": "constraint", "label": "Constraint" }
]
}
},
"initialSession": {
"promptA": "Use the context to answer the input.\n\nContext:\n{{context}}\n\nInput:\n{{input}}\n\nConstraint:\n{{constraint}}",
"cases": [
{
"input": "",
"expectedAnswer": "",
"variables": {
"context": "",
"constraint": ""
}
}
]
}
}With dynamicFields, placeholders such as {{context}}, {{input}}, and {{constraint}} automatically become labeled fields in each case. Learners edit the template as normal text, while placeholder values can differ across cases. Text written directly in the template stays shared. Expected output is used only for scoring, and each case can show its exact rendered prompt.
Work-in-progress (prompts, cases, settings, and the last results) is stored in eval-session.json. That file is local and not checked in. A saved session wins over initialSession on reload.
After each evaluation run, the server appends a numbered Evaluation section to .codesignal/report.md. The report keeps all evaluations from the current workspace, including each evaluation's setup, overall scores, cases, and individual runs. That file is gitignored.
npm run devnpm test- Node.js + Express
@anthropic-ai/sdk(Claude Messages API),openai(Chat Completions, including DeepSeek), or@google/genai(Gemini)- CodeSignal Bespoke Design System (git submodule)
- Vanilla JS + esbuild