Skip to content
Merged
48 changes: 24 additions & 24 deletions .env.example
Original file line number Diff line number Diff line change
@@ -1,37 +1,37 @@
# flux environment variables — copy to .env and fill in

# Provider API keys (or store them in the OS keychain — see
# docs/guides/CREDENTIAL-SETUP-FLOW.md)
FLUX_API_KEY=
OPENAI_API_KEY=
# Provider credentials, one per registry entry in catalog/registry/providers.go
# (registry SortOrder). Or store them in the OS keychain — see
# docs/guides/CREDENTIAL-SETUP-FLOW.md. catalog/registry/docs_test.go keeps this
# block in sync with the registry.
AGNES_API_KEY=
AWS_SECRET_ACCESS_KEY=
ANTHROPIC_API_KEY=
GEMINI_API_KEY=
AZURE_OPENAI_API_KEY=
CANOPYWAVE_API_KEY=
CLINE_API_KEY=
CONCENTRATE_API_KEY=
DEEPSEEK_API_KEY=
GEMINI_API_KEY=
GROQ_API_KEY=
KIMI_API_KEY=
MOONSHOT_API_KEY=
ZAI_API_KEY=
ZAI_CODING_API_KEY=
XIAOMI_MIMO_PAYG_API_KEY=
XIAOMI_MIMO_TOKEN_PLAN_API_KEY=
MINIMAX_API_KEY=
LONGCAT_API_KEY=
MINIMAX_PAYG_API_KEY=
MINIMAX_TOKEN_PLAN_API_KEY=
AZURE_OPENAI_API_KEY=
AWS_SECRET_ACCESS_KEY=
VERTEX_ACCESS_TOKEN=
OPENAI_API_KEY=
OPENCODEGO_API_KEY=
OPENROUTER_API_KEY=
CONCENTRATE_API_KEY=
OPENGATEWAY_API_KEY=
STEPFUN_API_KEY=
AGNES_API_KEY=
LONGCAT_API_KEY=
FIREWORKS_API_KEY=
CANOPYWAVE_API_KEY=
OLLAMA_BASE_URL=
POOLSIDE_API_KEY=
CLINE_API_KEY=
OPENCODEGO_API_KEY=
VERTEX_ACCESS_TOKEN=
XAI_API_KEY=
OLLAMA_BASE_URL=
XIAOMI_MIMO_PAYG_API_KEY=
XIAOMI_MIMO_TOKEN_PLAN_API_KEY=
ZAI_CODING_API_KEY=
ZAI_API_KEY=
STEP_API_KEY=
OPENGATEWAY_API_KEY=
FIREWORKS_API_KEY=

# Default model overrides
OPENAI_MODEL=gpt-4o
Expand Down
19 changes: 13 additions & 6 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,7 +8,8 @@ When starting any new work (feature, fix, refactor, chore), always create a feat

## Design Principles

- **Model-agnostic** — single interface for 75+ LLM providers
- **Model-agnostic** — single interface for the 28 provider gateways in
`catalog/registry/providers.go` (see README "Supported Providers")
- **Host-neutral engine** — Flux owns provider routing, transport, caching,
retry/fallback, and normalized telemetry; hosts own product UX and semantics
- **Streaming-first** — all responses are streamed; blocking is opt-in
Expand Down Expand Up @@ -55,6 +56,12 @@ make ci # Full CI suite
`StreamResult`, `ResponseFormat`, `ImageURLPart`, `InputAudioPart`) live in
`llm` with no `engine` alias; widening the facade to cover them is a
deliberate API change, not an incidental one.
- The facade already exposes a frozen set of engine-internal symbols
(`credentials.Store`/`MapStore`, `operationsgraph.Input`/`Export`, the
`provider/resilience` rate-limit config and `provider/cache.CacheConfig`),
listed in `docs/architecture/HOST-ENGINE-BOUNDARY.md`. Changing them breaks
hosts. `engine/host_surface_test.go` fails whenever that reachable set
changes; update its list and the doc together, deliberately.
- `provider/core.Provider` is the lower-level provider contract; keep its
method set stable and use it across feature packages
- Streaming tests need careful goroutine management
Expand All @@ -77,7 +84,7 @@ make ci # Full CI suite
- **Provider interface**: `provider/core.Provider` with `Chat()`, `StreamChat()`, `Ping()`, `Name()`
- **Core request types**: `provider/core.FluxMessage`, `FluxResponse`, `FluxTool`, `FluxUsage`
- **Config struct**: `provider/core.FluxConfig` with `Provider`, `APIKey`, `BaseURL`, `Model`, `MaxRetries`
- **Provider implementations**: `provider/adapters/AnthropicClient`, `OpenAIClient`, `GeminiClient`, etc.
- **Provider implementations**: `provider/adapters/anthropic.go` (`AnthropicClient`), `openai.go` (`OpenAIClient`), `gemini.go` (`GeminiClient`), etc.
- **Compatibility configs**: `provider/adapters.OpenAICompat`, `GrokCompat`, `OpenRouterCompat`
- **Error type**: `FluxError` with `Provider`, `Op`, `StatusCode`, `RequestID`, `Message`, `Err` fields
- **Stream types**: `StreamResult`, `SSEEvent`, `StreamEvent` — streaming is SSE-based
Expand Down Expand Up @@ -140,14 +147,14 @@ make ci # Full CI suite
| Azure provider | `provider/adapters/azure.go` |
| Provider registry | `provider/adapters/provider_registry.go` |
| Provider compatibility | `provider/adapters/compat.go` (`OpenAICompat`, `GrokCompat`, etc.) |
| SSE streaming | `provider/stream.go` (`parseSSEStream()`, `SSEEvent`) |
| SSE streaming | `provider/core/stream.go` (`parseSSEStream()`, `SSEEvent`) |
| Retry logic | `provider/core/retry.go` (`RetryConfig`, `backoffDelay()`, `shouldRetry()`) |
| Rate limiting | `provider/resilience/ratelimit.go`, `provider/resilience/adaptive_ratelimit.go` |
| Caching | `provider/cache/cache.go`, `provider/cache/semantic_cache.go` |
| Fallback chains | `provider/resilience/fallback.go` |
| Fallback chains | `router/router.go` (fallback providers), `router/deployment_router.go` (fallback deployment stages) |
| Auto-continuation | `provider/resilience/continuation.go` |
| Error types | `provider/errors.go` (`FluxError`, `IsRetriable()`, `IsAuthError()`) |
| Error constants | `errors/errors.go` (API error messages, prompt-too-long parsing) |
| Error types | `provider/core/errors.go` (`FluxError`, `IsRetriable()`, `IsAuthError()`) |
| Error constants | `types/errors.go` (API error messages, prompt-too-long parsing) |
| Model catalog | `catalog/` (pricing, context windows, capabilities per provider) |
| Credentials | `credentials/` (key storage, env detection, scrubbing) — `HasSecret` is silent on miss (boolean predicate); `LookupSecret` logs `Debug` on `ErrNotFound` and `Warn` on real backend errors |
| Mock provider | `provider/testkit/mock.go` |
Expand Down
122 changes: 73 additions & 49 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,6 +61,13 @@ Everything else is engine-internal: `provider`, `catalog`, `config`,
contracts. Enforced by `rho/scripts/check-flux-engine-boundary.sh`
and two Go AST tests in `rho/internal/testaudit/`.

Exception: a fixed set of engine-internal symbols is reachable through the
facade (for example `credentials.Store` behind `engine.Options.SecretStore`
and `operationsgraph.Input` behind `engine.OperationsGraphInput`). Those
symbols are frozen as part of the contract; see
[Frozen engine-internal types](docs/architecture/HOST-ENGINE-BOUNDARY.md#frozen-engine-internal-types).
`engine/host_surface_test.go` fails when that set changes.

- do not import `rho/internal/*`
- do not import the removed legacy path `rho/shared/types`

Expand All @@ -70,8 +77,9 @@ and two Go AST tests in `rho/internal/testaudit/`.
go get github.com/GrayCodeAI/flux
```

Requires Go 1.26+ and a configured provider credential. Minimal dependencies
(UUID, OpenTelemetry, SQLite, keyring).
Requires Go 1.26+ and a configured provider credential. Direct dependencies:
UUID, tiktoken tokenizer, OS keyring, OpenTelemetry, pure-Go SQLite, and gRPC
(linked only into `-tags grpc` builds of `internal/grpc`).

```go
import (
Expand Down Expand Up @@ -171,9 +179,9 @@ Named `primary` / `weak` / `editor` model slots with fallback to primary, plus a

`POST /rerank` endpoint (provider-backed with lexical fallback) and a `GET /ready` readiness probe alongside the existing health check.

### gRPC Skeleton
### gRPC Transport (opt-in, internal)

Dependency-free gRPC API skeleton behind the `grpc` build tag — wired when generated stubs are available.
`internal/grpc` holds an optional gRPC transport behind the `grpc` build tag. It serves `flux.v1.ChatService/Chat` with a registered `json` content subtype (no `.proto` files or generated stubs; clients call with `grpc.CallContentSubtype("json")`), backed by `EngineChatService` over `conversation.Engine`. The package is internal, so hosts cannot import it, and nothing in flux starts it. `google.golang.org/grpc` is a direct requirement in `go.mod`, so it appears in consumers' module graphs, but only `-tags grpc` builds link it.

## Documentation

Expand All @@ -199,40 +207,40 @@ ANTHROPIC_API_KEY=sk-... go run ./examples/basic/

## Supported Providers

28 provider gateways in `catalog/registry/providers.go` (rho `/config` uses the same list), listed in registry `SortOrder`:
28 provider gateways in `catalog/registry/providers.go` (rho `/config` uses the same list), listed in registry `SortOrder`. `catalog/registry/docs_test.go` fails when this table, the count, or `.env.example` drift from the registry.

| Provider | ID | Env variable |
|---|---|---|
| **Agnes** | `agnes` | `AGNES_API_KEY` |
| **Amazon Bedrock** | `bedrock` | `AWS_SECRET_ACCESS_KEY` (+ `AWS_ACCESS_KEY_ID`, `AWS_SESSION_TOKEN`) |
| **Anthropic** | `anthropic` | `ANTHROPIC_API_KEY` |
| **OpenAI** | `openai` | `OPENAI_API_KEY` |
| **Google Gemini** | `gemini` | `GEMINI_API_KEY` |
| **DeepSeek** | `deepseek` | `DEEPSEEK_API_KEY` |
| **xAI (Grok)** | `grok` | `XAI_API_KEY` |
| **Kimi (Moonshot)** | `kimi` | `MOONSHOT_API_KEY` |
| **Z.AI — Coding Plan** | `zai_coding` | `ZAI_CODING_API_KEY` |
| **Z.AI — Pay-as-you-go** | `zai_payg` | `ZAI_API_KEY` |
| **Xiaomi (MiMo) Token Plan** | `xiaomi_mimo_token_plan` | `XIAOMI_MIMO_TOKEN_PLAN_API_KEY` (+ region `cn` / `sgp` / `ams`) |
| **Xiaomi (MiMo) Pay-as-you-go** | `xiaomi_mimo_payg` | `XIAOMI_MIMO_PAYG_API_KEY` |
| **MiniMax — Token Plan** | `minimax_token_plan` | `MINIMAX_TOKEN_PLAN_API_KEY` |
| **MiniMax — Pay-as-you-go** | `minimax_payg` | `MINIMAX_PAYG_API_KEY` |
| **Azure OpenAI** | `azure` | `AZURE_OPENAI_API_KEY` (+ `AZURE_OPENAI_ENDPOINT`) |
| **Amazon Bedrock** | `bedrock` | `AWS_SECRET_ACCESS_KEY` (+ `AWS_ACCESS_KEY_ID`, `AWS_SESSION_TOKEN`) |
| **Vertex AI** | `vertex` | `VERTEX_ACCESS_TOKEN` (or `GOOGLE_OAUTH_ACCESS_TOKEN`) |
| **OpenRouter** | `openrouter` | `OPENROUTER_API_KEY` |
| **CanopyWave** | `canopywave` | `CANOPYWAVE_API_KEY` |
| **Poolside** | `poolside` | `POOLSIDE_API_KEY` |
| **Groq** | `groq` | `GROQ_API_KEY` |
| **ClinePass** | `clinepass` | `CLINE_API_KEY` |
| **Concentrate** | `concentrate` | `CONCENTRATE_API_KEY` |
| **OpenGateway** | `opengateway` | `OPENGATEWAY_API_KEY` |
| **StepFun** | `stepfun` | `STEPFUN_API_KEY` |
| **Agnes** | `agnes` | `AGNES_API_KEY` |
| **DeepSeek** | `deepseek` | `DEEPSEEK_API_KEY` |
| **Google Gemini** | `gemini` | `GEMINI_API_KEY` |
| **Groq** | `groq` | `GROQ_API_KEY` |
| **Kimi (Moonshot)** | `kimi` | `MOONSHOT_API_KEY` |
| **LongCat** | `longcat` | `LONGCAT_API_KEY` |
| **Fireworks AI** | `fireworks` | `FIREWORKS_API_KEY` |
| **MiniMax — Pay-as-you-go** | `minimax_payg` | `MINIMAX_PAYG_API_KEY` |
| **MiniMax — Token Plan** | `minimax_token_plan` | `MINIMAX_TOKEN_PLAN_API_KEY` |
| **OpenAI** | `openai` | `OPENAI_API_KEY` |
| **OpenCode Go** | `opencodego` | `OPENCODEGO_API_KEY` |
| **OpenRouter** | `openrouter` | `OPENROUTER_API_KEY` |
| **Ollama** | `ollama` | `OLLAMA_BASE_URL` (local; no API key) |
| **Poolside** | `poolside` | `POOLSIDE_API_KEY` |
| **Vertex AI** | `vertex` | `VERTEX_ACCESS_TOKEN` (or `GOOGLE_OAUTH_ACCESS_TOKEN`) |
| **xAI (Grok)** | `grok` | `XAI_API_KEY` |
| **Xiaomi (MiMo) Pay-as-you-go** | `xiaomi_mimo_payg` | `XIAOMI_MIMO_PAYG_API_KEY` |
| **Xiaomi (MiMo) Token Plan** | `xiaomi_mimo_token_plan` | `XIAOMI_MIMO_TOKEN_PLAN_API_KEY` (+ region `cn` / `sgp` / `ams`) |
| **Z.AI — Coding Plan** | `zai_coding` | `ZAI_CODING_API_KEY` (+ region `international` / `cn`) |
| **Z.AI — Pay-as-you-go** | `zai_payg` | `ZAI_API_KEY` (+ region `international` / `cn`) |
| **StepFun** | `stepfun` | `STEP_API_KEY` (+ region `global` / `cn`) |
| **OpenGateway** | `opengateway` | `OPENGATEWAY_API_KEY` |
| **Fireworks AI** | `fireworks` | `FIREWORKS_API_KEY` |

Runtime auto-detection uses a separate priority order for chat when no deployment is pinned; see `config` profiles.
Runtime auto-detection uses a separate priority order (`config.APIProviderDetectionOrder`) when no deployment is pinned.

## Usage

Expand Down Expand Up @@ -289,38 +297,54 @@ config.SaveProviderConfig(cfg, "") // save changes

```
flux/
├── engine/ # Stable host-facing facade and provider-neutral DTOs
├── provider/ # Provider runtime and feature packages
├── engine/ # Stable host-facing facade (hosts import engine, llm, graph, tools)
├── llm/ # Host-facing DTOs and the Provider port that engine re-exports
├── graph/ # Portable execution-graph vocabulary
├── tools/ # Tool-call and tool-result contracts
├── provider/ # Provider runtime composition root (FluxClient)
│ ├── core/ # Provider-neutral wire, stream, retry, and transport primitives
│ ├── adapters/ # Provider protocol adapters and construction registry
│ └── embeddings/ # Embedding clients, cache, and defaults
├── config/ # Provider configuration & routing
│ └── credential/ # Credential file management
│ ├── resilience/ # Rate limits, continuation, guardrails, and error policy
│ ├── cache/ # Response and semantic caches
│ ├── batch/ # Batch execution
│ ├── embeddings/ # Embedding clients, cache, and defaults
│ ├── media/ # Image and audio clients, structured prompts
│ ├── extraction/ # Structured extraction
│ ├── observability/ # Usage, cost, metrics, tracing, and recording
│ └── testkit/ # Mock provider for tests
├── catalog/ # Model catalog & tier system
│ ├── registry/ # Provider registry (single source of truth for providers)
│ ├── discover/ # Model discovery
│ ├── legacy/ # Legacy model support
│ ├── live/ # Live model data
│ └── registry/ # Model registry
├── codeagent/ # Code agent retry & fallback strategies
├── conversation/ # Conversation engine with branching
├── credentials/ # Credential management
├── docs/ # Documentation & guides
├── examples/ # Runnable code examples
├── router/ # Provider routing strategies
│ ├── live/ # Live model listing per provider
│ ├── capabilities/ # Capability and deprecation data
│ └── concentrate/ opencodego/ opengateway/ xiaomi/ zai/ # Gateway-specific helpers
├── config/ # Provider configuration & routing
│ └── credential/ # Credential file management
├── credentials/ # Keyring/env credential stores and OIDC keyless auth
├── router/ # Routing strategies, deployment router, circuit breakers
│ └── controlplane/ # Versioned, signed peer manifests and replicas
├── runtime/ # Engine-internal provider/model/credential resolution
├── setup/ # Catalog-backed deployment wiring
├── operationsgraph/ # Privacy-safe route and generation telemetry projection
├── runtime/ # Runtime manifest & routing policies
├── storage/ # SQLite conversation DAG store
├── types/ # Branded types & API errors
├── errors/ # Error message constants
├── conversation/ # Conversation engine with branching
├── storage/ # SQLite conversation DAG store, virtual keys, budgets
├── codeagent/ # Code agent retry & fallback strategies
├── verify/ # Provider conformance harness
├── types/ # Shared message types & API errors
├── constants/ # API limits
├── utils/ # Error utilities
├── api/ # OpenAPI spec for internal/api
├── internal/
│ ├── api/ # HTTP API handlers
│ ├── cache/ # Response cache warmer
│ ├── api/ # HTTP API server (library code; no flux binary starts it)
│ ├── cache/ # Cache backends and response cache warmer
│ ├── grpc/ # Optional gRPC transport (build tag grpc)
│ ├── health/ # Provider health checker
│ ├── observability/ # OpenTelemetry spans & metrics
│ ├── sdk/ # Go, Python, TypeScript client SDKs
│ └── version/ # Version information
│ ├── httputil/ probehttp/ shrink/ # HTTP, probe, and tool-description helpers
│ ├── observability/ # OpenTelemetry spans, metrics, and audit sinks
│ └── sdk/ # Go, Python, TypeScript clients for the internal/api HTTP surface
├── docs/ # Documentation & guides
├── examples/ # Runnable code examples
├── scripts/ # CI guards and helper scripts
└── assets/ # Logo and branding
```

Expand Down
Loading
Loading