SCAL-336134: Add developer examples related to chat history for spotter mcp server - #67
Open
mouryabalabhadra wants to merge 4 commits into
Open
mouryabalabhadra wants to merge 4 commits into
mouryabalabhadra wants to merge 4 commits into
Conversation
…er mcp server
Brings the `python-react-agent-simple-ui` example up to date with the Spotter 3 MCP toolset, and fixes the things that made it unreliable to run: embed auth, chart rendering after an answer expires, CSP-blocked iframes, and an MCP connection that failed intermittently and was slow to open. It also makes the agent faster, converts the client to TypeScript, and adds a README guide for integrating the ThoughtSpot MCP server into your own application.
- `claude_agent_mcp_server_v2.py` → `claude_agent_with_spotter3_mcp_server.py`
- `claude_agent_mcp_server_v2_with_chat_history.py` → `claude_agent_with_spotter3_mcp_server_and_chat_history.py`
- README and `env.template` re-worded off `v1`/`v2` onto Spotter 3 terms, and the OpenAI / Azure OpenAI (v1) documentation removed.
The client's `getAuthToken` returned one constant `VITE_TS_AUTH_TOKEN`. The SDK requires a *fresh* token per call: once that static token stops verifying, the SDK reports a duplicate token and the callback can never recover, because it hands back the same string.
- New `GET /api/ts-token` on both servers, minting a short-lived token per request via `POST /api/rest/2.0/auth/token/full`.
The server used the static `TS_AUTH_TOKEN` for its own MCP and REST calls, so every tool call failed with "Failed to validate connection" once that token expired, while the tool *list* still loaded.
- `server_token()` mints the server's own token and caches it until shortly before expiry. Its lifetime is set by `TS_SERVER_TOKEN_VALIDITY_SEC` (default 3600).
- Expiry is tracked with `time.time()`, not `time.monotonic()`: on macOS the monotonic clock pauses during sleep, which would keep an expired token in use after a laptop sleep.
- A background task (`keep_server_token_fresh()`) mints the token at startup and renews it about 12 minutes before expiry (at half-life for short lifetimes). Requests holding a valid token never wait on a renewal in progress. Previously a request that found the token expired paid the whole mint, which took ~40s on a staging cluster.
- A static `TS_AUTH_TOKEN` is now needed only when minting credentials are not set. The server exits with a clear message when neither is configured.
A ThoughtSpot answer object lives ~8 hours. The chat history stored each answer's `iframe_url` and `answer_id` (which is really a `{session_id, gen_no}` pair), so reopening an older conversation rendered a row of dead embeds.
- The server no longer persists those fields, and strips them on read too, so rows written before this change take the same path.
- `GET /api/conversations/{id}` now returns each answer's `answer_index` plus the conversation's `analytical_session_id`; the client emits a resolver placeholder and the Visual Embed SDK resolves a live URL. **Requires the SDK change in PR 2.**
- `reconcile_answers` asks `getConversation` how many answers each turn actually has, rather than trusting the stored copy. This also recovers answers from turns whose SSE stream was cut off mid-flight: the Agent finishes regardless, so the answer exists on ThoughtSpot's side even when the app recorded none of it.
`TS_MCP_API_VERSION` now defaults to `latest`, and the update readers accept both the server's digested shape and the Agent's raw shape (`text-chunk` vs `text_chunk`, `metadata.type == "thinking"` vs `is_thinking`, …), so `&enable-raw-session-updates=true` can be turned on through `TS_MCP_URL` without a code change. Intermediate "thinking" answers are filtered out — one three-question chat produced eleven answer updates but a single settled one — which also keeps our answer ordering aligned with `getConversation`'s.
A timed turn ("total sales in 2023") split into: MCP connection setup 9.7s, four Claude calls 7.6s, ThoughtSpot tool calls 51.6s.
- **Shared MCP session.** Opening a session costs several round trips (discovery probe, `initialize`, `tools/list`), measured at 6-11s and previously paid on every chat message. One session is now shared across turns (`McpPool` / `McpTurn`). Measured: connection setup went from 8.3s on the first message to 0.0s on the second; the turn went from 66.9s to 49.2s.
- The session is opened and closed by its own background task (anyio requires the task that opened a session to close it).
- It is replaced on token rotation; the old session stays open until the turns using it finish. It is also replaced when it breaks.
- A session replaced mid-turn stays open until that turn ends, so a parallel call still running on it is not cut off.
- **Retries.**
- A failed handshake is retried once with a freshly minted token. `initialize` intermittently returns a bare HTTP 500, which surfaced in the UI as `MCPError: Server returned an error response`.
- A call that fails with `-32600 "Session terminated"` is retried once on a fresh session. The check matches the message as well as the code, because the MCP client also uses -32600 when a call may already have run.
- **No spurious error after each answer.** The SSE stream no longer cancels the agent task after `done`, and a browser hang-up is treated as a cancel, not an error. In the history server, a failed SQLite save on an error path is logged rather than raised, so the browser always receives its `error` event.
- Default model is now `claude-haiku-4-5` (overridable with `ANTHROPIC_MODEL`). The agent only orchestrates ThoughtSpot's tools; measured on one question, its four Claude calls took ~4s on Haiku 4.5, ~8s on Opus 5 and ~17s on Sonnet 5.
- `model_request_options()` sends adaptive thinking only to models that support it; Haiku 4.5 runs without thinking.
- The server-side refusal fallback is sent only to models with fallback targets. `/v1/models` reports none for Sonnet 5 or Haiku 4.5.
- `[Timing]` logs in the chat-history server break each turn down into MCP session, Claude calls, tool calls and total.
- **TypeScript.** `App.jsx` → `App.tsx` and `main.jsx` → `main.tsx`, with typed messages, answers and SSE events. Adds `tsconfig.json` (type-check only; Vite still builds), `vite-env.d.ts`, and `npm run typecheck`.
- **Loading state for stored chats.** Opening a stored chat shows a spinner and marks the sidebar entry busy until it loads (2.5-5s in testing, while answers are reconciled against ThoughtSpot).
- Stale responses are ignored if another chat is opened or a new one started meanwhile.
- Sending is disabled while a chat loads, so a message can't go to the previous conversation.
- **Status line** shows the Analytics Agent's steps ("Searching for Datasets") instead of its reasoning text streamed one word-sized fragment at a time.
- **System theme.** `App.css` moved fully onto CSS custom properties with a `prefers-color-scheme: dark` block (no hardcoded colours left bypassing the tokens), and the embed itself gets matching dark `customizations` variables so the chart doesn't stay white inside a dark page.
- **`AnswerFrame` removed.** Answer iframes are injected as markup so React owns only the wrapper and doesn't fight the renderer's `replaceWith()`.
- **README: "Integrating the ThoughtSpot MCP server into your own application".** An eight-step guide, each step naming the function to copy:
1. ThoughtSpot and Anthropic prerequisites
2. Minting tokens on the server
3. Connecting to MCP
4. Giving the tools to the model
5. Streaming to the UI
6. Rendering charts with the Visual Embed SDK
7. Optional chat history
8. A production checklist
- **README fixes.**
- Ports: the frontend is `:8000` (strict, because it is the CSP-allowlisted origin) and the backend `:8001`. The `uvicorn` commands now pass `--port 8001`; previously they clashed with Vite on port 8000.
- Token-minting environment variables.
- Snippets updated to `App.tsx`.
- How stored charts are replayed.
- The shared MCP session.
- New troubleshooting rows.
- **`env.template`:** documents the model default and `TS_SERVER_TOKEN_VALIDITY_SEC`, and says when the static token is required.
- Dependency bumps: `anthropic>=1.2.0,<2`, `mcp>=2.1.1,<3`, `httpx` → `httpx2`, FastAPI/uvicorn; `@thoughtspot/visual-embed-sdk` `1.45.3-mcp.2` → `^1.52.1`; client dev dependencies `typescript`, `tslib`, `@types/react`, `@types/react-dom`.
- `.gitignore`: local `*.db` / `-wal` / `-shm` chat-history files.
- Live chats through both servers: answers stream, embeds render, follow-up turns reuse the same analytical session. Timed turns back to back confirm the second one skips MCP connection setup.
- `/api/ts-token` verified minting: `minted: true`, a different token per call, and the minted token authenticates against `/callosum/v1/session/isactive`.
- `/api/conversations` list/open/delete exercised against the stored database; the loading indicator and disabled input verified in the browser.
- Live MCP session tests:
- Ending the session out-of-band triggers one reconnect, and the call succeeds.
- On token rotation, the old session stays open for its turn, then closes.
- Hanging up mid-turn and mid-handshake leaves no traceback, and the next request works.
- Unit tests with fakes:
- A -32600 with a different message is not retried.
- A parallel call on a replaced session finishes normally.
- Background renewal: minting at startup; a request during a slow renewal returns the current token in 0ms; failed renewals keep the old token and retry; a short token life renews at half-life.
- Startup with minting only (no static token) succeeds; with neither it fails with a clear message.
- `py_compile` on both servers; `tsc` on the client (also passes with `--strict`); client `vite build` passes.
- A separate review agent audited the diff; its findings are fixed in this PR.
- `server/agent.py` (the OpenAI/Azure v1 backend) and the `openai` entry in `requirements.txt` are left in place but are no longer documented. Say if they should be removed.
- The client renders model markdown with `rehypeRaw` without sanitizing it, and does not validate iframe `src` values. Both are listed in the README's production checklist.
- On Haiku 4.5 the prompt cache does not engage: the ~3.6K-token prefix is under its 4,096-token minimum.
mouryabalabhadra
force-pushed
the
SCAL-336134
branch
from
September 28, 2026 09:17
8f9a473 to
7739fd9
Compare
ashubham
reviewed
Sep 30, 2026
| @@ -0,0 +1,1104 @@ | |||
| """ | |||
Member
There was a problem hiding this comment.
The servers seem to be too verbose, can they be made simpler?
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Brings the
python-react-agent-simple-uiexample up to date with the Spotter 3 MCP toolset, and fixes the things that made it unreliable to run: embed auth, chart rendering after an answer expires, CSP-blocked iframes, and an MCP connection that failed intermittently and was slow to open. It also makes the agent faster, converts the client to TypeScript, and adds a README guide for integrating the ThoughtSpot MCP server into your own application.claude_agent_mcp_server_v2.py→claude_agent_with_spotter3_mcp_server.pyclaude_agent_mcp_server_v2_with_chat_history.py→claude_agent_with_spotter3_mcp_server_and_chat_history.pyenv.templatere-worded offv1/v2onto Spotter 3 terms, and the OpenAI / Azure OpenAI (v1) documentation removed.The client's
getAuthTokenreturned one constantVITE_TS_AUTH_TOKEN. The SDK requires a fresh token per call: once that static token stops verifying, the SDK reports a duplicate token and the callback can never recover, because it hands back the same string.GET /api/ts-tokenon both servers, minting a short-lived token per request viaPOST /api/rest/2.0/auth/token/full.The server used the static
TS_AUTH_TOKENfor its own MCP and REST calls, so every tool call failed with "Failed to validate connection" once that token expired, while the tool list still loaded.server_token()mints the server's own token and caches it until shortly before expiry. Its lifetime is set byTS_SERVER_TOKEN_VALIDITY_SEC(default 3600).time.time(), nottime.monotonic(): on macOS the monotonic clock pauses during sleep, which would keep an expired token in use after a laptop sleep.keep_server_token_fresh()) mints the token at startup and renews it about 12 minutes before expiry (at half-life for short lifetimes). Requests holding a valid token never wait on a renewal in progress. Previously a request that found the token expired paid the whole mint, which took ~40s on a staging cluster.TS_AUTH_TOKENis now needed only when minting credentials are not set. The server exits with a clear message when neither is configured.A ThoughtSpot answer object lives ~8 hours. The chat history stored each answer's
iframe_urlandanswer_id(which is really a{session_id, gen_no}pair), so reopening an older conversation rendered a row of dead embeds.GET /api/conversations/{id}now returns each answer'sanswer_indexplus the conversation'sanalytical_session_id; the client emits a resolver placeholder and the Visual Embed SDK resolves a live URL. Requires the SDK change in PR 2.reconcile_answersasksgetConversationhow many answers each turn actually has, rather than trusting the stored copy. This also recovers answers from turns whose SSE stream was cut off mid-flight: the Agent finishes regardless, so the answer exists on ThoughtSpot's side even when the app recorded none of it.TS_MCP_API_VERSIONnow defaults tolatest, and the update readers accept both the server's digested shape and the Agent's raw shape (text-chunkvstext_chunk,metadata.type == "thinking"vsis_thinking, …), so&enable-raw-session-updates=truecan be turned on throughTS_MCP_URLwithout a code change. Intermediate "thinking" answers are filtered out — one three-question chat produced eleven answer updates but a single settled one — which also keeps our answer ordering aligned withgetConversation's.A timed turn ("total sales in 2023") split into: MCP connection setup 9.7s, four Claude calls 7.6s, ThoughtSpot tool calls 51.6s.
Shared MCP session. Opening a session costs several round trips (discovery probe,
initialize,tools/list), measured at 6-11s and previously paid on every chat message. One session is now shared across turns (McpPool/McpTurn). Measured: connection setup went from 8.3s on the first message to 0.0s on the second; the turn went from 66.9s to 49.2s.Retries.
initializeintermittently returns a bare HTTP 500, which surfaced in the UI asMCPError: Server returned an error response.-32600 "Session terminated"is retried once on a fresh session. The check matches the message as well as the code, because the MCP client also uses -32600 when a call may already have run.No spurious error after each answer. The SSE stream no longer cancels the agent task after
done, and a browser hang-up is treated as a cancel, not an error. In the history server, a failed SQLite save on an error path is logged rather than raised, so the browser always receives itserrorevent.Default model is now
claude-haiku-4-5(overridable withANTHROPIC_MODEL). The agent only orchestrates ThoughtSpot's tools; measured on one question, its four Claude calls took ~4s on Haiku 4.5, ~8s on Opus 5 and ~17s on Sonnet 5.model_request_options()sends adaptive thinking only to models that support it; Haiku 4.5 runs without thinking.The server-side refusal fallback is sent only to models with fallback targets.
/v1/modelsreports none for Sonnet 5 or Haiku 4.5.[Timing]logs in the chat-history server break each turn down into MCP session, Claude calls, tool calls and total.TypeScript.
App.jsx→App.tsxandmain.jsx→main.tsx, with typed messages, answers and SSE events. Addstsconfig.json(type-check only; Vite still builds),vite-env.d.ts, andnpm run typecheck.Loading state for stored chats. Opening a stored chat shows a spinner and marks the sidebar entry busy until it loads (2.5-5s in testing, while answers are reconciled against ThoughtSpot).
Status line shows the Analytics Agent's steps ("Searching for Datasets") instead of its reasoning text streamed one word-sized fragment at a time.
System theme.
App.cssmoved fully onto CSS custom properties with aprefers-color-scheme: darkblock (no hardcoded colours left bypassing the tokens), and the embed itself gets matching darkcustomizationsvariables so the chart doesn't stay white inside a dark page.AnswerFrameremoved. Answer iframes are injected as markup so React owns only the wrapper and doesn't fight the renderer'sreplaceWith().README: "Integrating the ThoughtSpot MCP server into your own application". An eight-step guide, each step naming the function to copy:
README fixes.
:8000(strict, because it is the CSP-allowlisted origin) and the backend:8001. Theuvicorncommands now pass--port 8001; previously they clashed with Vite on port 8000.App.tsx.env.template: documents the model default andTS_SERVER_TOKEN_VALIDITY_SEC, and says when the static token is required.Dependency bumps:
anthropic>=1.2.0,<2,mcp>=2.1.1,<3,httpx→httpx2, FastAPI/uvicorn;@thoughtspot/visual-embed-sdk1.45.3-mcp.2→^1.52.1; client dev dependenciestypescript,tslib,@types/react,@types/react-dom..gitignore: local*.db/-wal/-shmchat-history files.Live chats through both servers: answers stream, embeds render, follow-up turns reuse the same analytical session. Timed turns back to back confirm the second one skips MCP connection setup.
/api/ts-tokenverified minting:minted: true, a different token per call, and the minted token authenticates against/callosum/v1/session/isactive./api/conversationslist/open/delete exercised against the stored database; the loading indicator and disabled input verified in the browser.Live MCP session tests:
Unit tests with fakes:
Startup with minting only (no static token) succeeds; with neither it fails with a clear message.
py_compileon both servers;tscon the client (also passes with--strict); clientvite buildpasses.A separate review agent audited the diff; its findings are fixed in this PR.
server/agent.py(the OpenAI/Azure v1 backend) and theopenaientry inrequirements.txtare left in place but are no longer documented. Say if they should be removed.