Skip to content

Repository files navigation

Prompt Evaluation Simulator (learn_cosmo-prompteval)

A CodeSignal learning project that teaches how to evaluate LLM prompts — not only how to write them.

Learners run the same prompt template against an input across multiple independent LLM calls, then score outputs against an optional expected answer using simple metrics.

Milestone status

  1. Done — Prompt template + input + 1–5 independent runs → collect outputs
  2. Done — Optional expected answer + metrics + mean/min/max
  3. Done — Prompt A vs Prompt B under shared conditions (winner by mean)
  4. Done — Evaluation across multiple test cases (overall + per-case scores)
  5. Next — Charts / distributions, or model/provider comparison

Setup

git clone --recurse-submodules <this-repo-url>
cd learn_cosmo-prompteval
npm install
cp .env.example .env
cp session.config.example.json session.config.json

Fill in .env with the API key (and optional *_BASE_URL) for the provider you want. Choose the model in session.config.json as provider/model-id:

  • anthropic/claude-sonnet-4-6 — needs ANTHROPIC_API_KEY, optional ANTHROPIC_BASE_URL
  • openai/gpt-5.6-luna — needs OPENAI_API_KEY, optional OPENAI_BASE_URL
  • google/gemini-3.6-flash — needs GOOGLE_API_KEY, optional GOOGLE_BASE_URL (gemini/… also routes to Gemini)
  • ~deepseek/deepseek-v4-flash-latest — needs DEEPSEEK_API_KEY and DEEPSEEK_BASE_URL (deepseek/… and deepseek-ai/… also route here; uses the OpenAI SDK). If both DeepSeek vars are unset, it reuses OPENAI_API_KEY / OPENAI_BASE_URL (production proxy hack).

session.config.json is separate from .env. It is local (not checked in) and holds session defaults, not secrets:

  • model (optional) — provider/model-id (default anthropic/claude-sonnet-4-6); must be listed in allowedModels
  • allowedModels (optional) — picker list of provider/model-id refs (defaults to Anthropic, OpenAI, Gemini, and DeepSeek examples above)
  • allowUserModelSelection (optional) — when true, show a model picker and let the saved eval session override model with an entry from allowedModels (default false)
  • allowCompare (optional) — when true, show “Compare with another prompt” so learners can A/B two prompts (default false)
  • allowedMetricIds (optional) — metrics shown in the picker and accepted by the API. When omitted, the original Course 1 metrics remain unchanged: exact-match, exact-match-ci, contains, string-similarity, and word-overlap-f1. Opt in to newer validation with regex-match, valid-json, and/or llm-judge.
  • llmJudgeModel (optional) — fixed provider/model-id used only by llm-judge. The server controls this value and the UI displays it to learners. When omitted, the generation model is reused for backward compatibility.
  • features.promptTemplating (optional) — enables reusable named blanks, shared examples, and a filled-in prompt preview. Missing configuration keeps the current Course 1 UI and {{input}} behavior unchanged.
  • maxConcurrency (optional) — max in-flight LLM calls during an evaluation (default 4, range 1–50). Set to 1 for serial.
  • defaults (optional) — runs sets the initial run count while minRuns, maxRuns, minCases, and maxCases set the editable limits (each 1–5)
  • initialSession (optional) — promptA, promptB, and cases (input / expectedAnswer)

Without session.config.json, prompts and cases start empty and the UI uses the built-in 1–5 limits. Copy session.config.example.json to prefill the capital-city demo.

For example, a later course can enable every validation type without changing Course 1:

{
  "llmJudgeModel": "anthropic/claude-sonnet-4-6",
  "allowedMetricIds": [
    "exact-match",
    "regex-match",
    "valid-json",
    "llm-judge"
  ]
}

regex-match treats Expected Answer as a regular expression. valid-json is a deterministic function checker and does not require an expected answer. llm-judge makes a second call to llmJudgeModel for each generated output and requires an expected answer. Function checkers are registered in code and enabled by ID; configuration never executes arbitrary JavaScript.

Config-gated prompt templating

Later-course tasks can let the prompt template define the fields shown in every case without changing Course 1:

{
  "allowCompare": false,
  "features": {
    "promptTemplating": {
      "enabled": true,
      "templateEditable": true,
      "showPreview": true,
      "dynamicFields": true,
      "fields": [
        { "name": "context", "label": "Context" },
        { "name": "input", "label": "Input" },
        { "name": "constraint", "label": "Constraint" }
      ]
    }
  },
  "initialSession": {
    "promptA": "Use the context to answer the input.\n\nContext:\n{{context}}\n\nInput:\n{{input}}\n\nConstraint:\n{{constraint}}",
    "cases": [
      {
        "input": "",
        "expectedAnswer": "",
        "variables": {
          "context": "",
          "constraint": ""
        }
      }
    ]
  }
}

With dynamicFields, placeholders such as {{context}}, {{input}}, and {{constraint}} automatically become labeled fields in each case. Learners edit the template as normal text, while placeholder values can differ across cases. Text written directly in the template stays shared. Expected output is used only for scoring, and each case can show its exact rendered prompt.

Work-in-progress (prompts, cases, settings, and the last results) is stored in eval-session.json. That file is local and not checked in. A saved session wins over initialSession on reload.

After each evaluation run, the server appends a numbered Evaluation section to .codesignal/report.md. The report keeps all evaluations from the current workspace, including each evaluation's setup, overall scores, cases, and individual runs. That file is gitignored.

Run

npm run dev

Open http://localhost:3000

Tests

npm test

Stack

  • Node.js + Express
  • @anthropic-ai/sdk (Claude Messages API), openai (Chat Completions, including DeepSeek), or @google/genai (Gemini)
  • CodeSignal Bespoke Design System (git submodule)
  • Vanilla JS + esbuild

About

Prompt Evaluation Simulator – evaluate LLM prompts with independent Octavus runs and metrics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages