Note for website readers: Links to source files work on GitHub and in VS Code.
They will not open on this website.
docs / FEATURES
Fred
Feature Reference
Concise, complete inventory of what the Fred platform provides — from document
ingestion and agent execution to platform governance and security certification.
What Fred does
Fred is a governed platform for deploying AI agents at team scale in regulated
environments. It covers the full lifecycle: ingesting knowledge, running agents
against that knowledge, and exposing the result through a managed chat interface
— all within strict team-scoped authorization boundaries.
Knowledge ingestion
32 file formats — documents, spreadsheets, images, audio, video — converted to
searchable, citation-rich Markdown via configurable processing profiles.
Agent execution
HTTP SSE streaming, HITL pause/resume, LangGraph checkpoints,
and a typed authorization envelope that keeps product concerns out of the execution engine.
MCP tool ecosystem
Agents discover and call tools via the Model Context Protocol.
Operators configure which servers each agent instance may use at enrollment time.
Platform governance
ReBAC (OpenFGA) enforces team scoping, role separation, and
resource quotas. Operators set immutable platform guardrails; teams configure within them.
Security certification
First homologation at C3 (French national security classification)
completed. Security posture is actively maintained and re-evaluated across platform updates.
Bring-your-own agents
Import fred-runtime and fred-sdk, define
agent behaviors, deploy as a container. The control plane enrolls custom pods
exactly like the bundled apps/fred-agents.
Services
| Service | Layer | Responsibility |
apps/control-plane-backend |
Control plane |
Teams, sessions, agent enrollment, execution authorization (JWT + per-request OpenFGA check), prompt library, admin APIs |
apps/knowledge-flow-backend |
Control plane |
Document ingestion pipeline, vector search, object storage, resource lifecycle |
apps/fred-agents |
Runtime |
Reference agent pod — ReAct agents, test harness, MCP tool integration |
apps/frontend |
Client |
React SPA — managed chat, agent catalog, knowledge management, session lifecycle |
libs/fred-runtime |
Runtime lib |
Agent execution framework — SSE, HITL, checkpoints, grant validation. Core import for custom agent pods. |
libs/fred-sdk |
SDK lib |
Typed execution contracts, UiPart primitives, prompt utilities, authoring helpers |
libs/fred-core |
Cross-cutting lib |
Caching, configuration, structured logging, KPI stores, team ID utilities |
Processing profiles
Each ingestion request selects a profile. Profiles are declared in
configuration.yaml under processing.profiles
and can be overridden per deployment.
| Profile | PDF engine | OCR | Tables | Images | Use for |
| fast default |
pypdfium2 |
No |
Preserve (no structure) |
No |
Text-native PDFs, high-volume ingestion |
| medium |
docling_parse |
OpenVINO |
Structure detection |
No |
Scanned PDFs, complex tables |
| rich |
docling_parse |
OpenVINO (full page) |
Structure + images |
Yes (captions) |
High-fidelity technical documents, figure extraction |
Text splitting (all profiles): chunk size 1500 tokens, overlap 150 tokens,
table preservation enabled. Configurable per deployment.
Audio & video transcription
AudioProcessor converts spoken content to timestamped Markdown transcripts.
Video files have their audio track extracted first; the intermediate WAV is deleted on
completion. The Whisper model is loaded lazily on first use.
| Property | Value |
| Engine | faster-whisper ≥ 1.1.0 |
| Video demux | PyAV ≥ 14.0.0 — resamples to 16 kHz mono WAV |
| Default model | base (configurable via audio_model.whisper_model_size) |
| Device | cpu default; cuda supported (configurable via audio_model.device) |
| Language | Auto-detected; overridable via audio_model.language |
| Output | Markdown with [HH:MM:SS] timestamp per segment, detected language, total duration |
| Beam size | 5 (Whisper default) |
Upload & folder management
The Resources (Corpus) workspace ingests whole folder trees, not just
individual files, and gives upfront feedback before spending upload
bandwidth or ingestion time.
| Feature | Detail |
| Full-page drag-and-drop |
The whole page is a drop target for the folder being viewed, not just a dropzone widget. A dropped directory keeps its on-disk structure: each subdirectory becomes a nested document tag, and files upload under their own subdirectory's tag instead of flattening into the drop target. |
| Subtree ingestion state |
Folder rows carry the same StatusChip status="processing" as document rows for as long as any document anywhere under them — own files plus every sub-folder, at any depth — is still ingesting, clearing once the last one settles. Unions the currently loaded tag view with the live SSE task feed. |
| Upload quota precheck |
POST /knowledge-flow/v1/quota/precheck checks a whole upload batch's declared total size against the destination's quota (owning team or personal space) in one round trip before any file bytes are sent. An over-quota batch is rejected wholesale; the drawer stays open showing the server's numbers, with Save disabled until the file list changes. |
| Folder-depth cap |
Dropped or manually created folder trees are capped at MAX_FOLDER_DEPTH = 15 levels, counting the destination path — enforced server-side in the tag-path validator (MAX_TAG_PATH_DEPTH, 422 past the cap) and pre-filtered client-side on every drop surface, with a toast naming the skipped file count. |
SSE event stream
Every agent execution produces a typed Server-Sent Events stream. Events are defined
in libs/fred-sdk and frozen — agents and frontends depend on this contract.
| Event | When | Key fields |
assistant_delta |
Each streaming text chunk |
delta: str |
status |
Generic operational progress signal (no chain-of-thought semantics) |
status: str, detail: str | None |
thought_start |
Opens a reasoning block |
thought_id, phase: ThoughtKind (planning · tool_use · observation · reflection · synthesis), title, source (authored | model_native) |
thought_delta |
Incremental text fragment into an open reasoning block |
thought_id, delta: str |
thought_end |
Closes a reasoning block |
thought_id, conclusion, duration_ms |
tool_call |
Agent invokes an MCP tool |
tool_name, call_id, arguments |
tool_result |
Tool returns result |
call_id, content, tool_name, is_error, sources[], ui_parts[], latency_ms |
final |
Turn complete |
content (full), sources: VectorSearchHit[], ui_parts[], token_usage, finish_reason |
awaiting_human |
HITL pause requested |
request: HumanInputRequest — stage, title, question, choices[], free_text, metadata, checkpoint_id |
node_error |
Graph node raised and declared on_error routing |
node_id, error_message, routed_to |
execution_error |
Unhandled pipeline exception (terminal) |
message: str |
UI rendering parts (UiPart)
Rich structured output emitted in the final event ui_parts[] array.
| Type | Fields | Frontend action |
LinkPart |
href, title, kind: "download" | "open" | "cite" |
Rendered as action button or inline citation |
GeoPart |
GeoJSON FeatureCollection |
Rendered as interactive map component |
Execution endpoints
| Endpoint | Mode |
POST /agents/execute/stream | SSE streaming (primary) |
POST /agents/execute | Synchronous request/response |
GET /agents/sessions/{session_id}/messages | Full turn history as ChatMessage[] |
HITL & checkpoints
Human-in-the-loop pause lets an agent request a decision or input mid-execution.
The graph state is persisted to a LangGraph checkpoint; the session resumes from
that exact point when the human responds.
| Capability | Detail |
| Pause trigger |
Agent emits awaiting_human event with a HumanInputRequest (stage, title, question, choices[], free_text, checkpoint_id) |
| Resume |
POST /agents/execute/stream with execution_action: "resume" + checkpoint_id + human response carried in ExecutionConfig.resume_payload |
| Persistence |
LangGraph checkpoint storage; HITL request and response persisted as hitl_request / hitl_response channels in session history |
| Stateless execution guarantee |
No in-memory session state between requests; full state in checkpoint only |
| CLI inspection |
/checkpoints [limit], /checkpoint <thread_id>, /stats |
Execution authorization
There is no ExecutionGrant and no control-plane-issued signed or
unsigned token (RFC RUNTIME-07 rev. 2 — the earlier signed-grant / valet-key
design was reversed by RFC decision D5). Identity is the caller's own Keycloak
JWT, forwarded on every request; each agent pod is its own execution authority
and authorizes the request itself. The frontend cannot forge or escalate access
because there is no capability token to forge — access is derived fresh from
the JWT and a live OpenFGA check on every call.
| Step | Detail |
| Authentication | Caller's Keycloak JWT in the Authorization: Bearer header; the pod is an OAuth2 resource server and validates issuer/audience under the c3 security profile |
| Audience binding | Each pod validates aud == its own client_id — per-agent audience, anti-confused-deputy (decision D5c) |
| Authorization | Per-request pod-side OpenFGA CAN_READ check on runtime_context.team_id; a missing or unauthorized team fails closed (403) |
| Identity integrity | user_id is taken from the validated token only — any body-supplied user id is neutralized |
| Execution action | execution_action: execute or resume (HITL) — the one field that survives from the retired ExecutionGrant envelope |
Prepare-execution flow: POST /teams/{team_id}/agent-instances/{id}/prepare-execution
returns ExecutionPreparation — execute_url, execute_stream_url,
messages_url_template, and capability flags (supports_streaming,
supports_hitl, supports_ui_parts) plus computed chat_controls[].
No grant object, no expiry — every subsequent call is authorized independently
by the runtime pod using the caller's own JWT.
Developer CLI (fred-agents-cli)
A REPL for interacting with any runtime agent directly — authentication, session
management, checkpoint inspection, and KPI display — without the frontend.
| Command | Description |
/agents, /agent <id> | List available agents and switch active agent |
/session <id>, /sessions | Set session scope or list existing sessions |
/history [session_id] | Print full message history for a session |
/checkpoints [limit], /checkpoint <id> | List or inspect LangGraph checkpoint state |
/stats | Checkpoint storage statistics |
/team [team_id|clear] | Control team scope for managed execution |
/context | Print active execution context summary |
/kpi [pattern] | Display runtime KPI counters and timings |
/login, /login-password | PKCE browser flow or password authentication |
/whoami, /logout | Auth status and session termination |
/help <question> | Ask the active agent a question in natural language |
Bundled agent templates
apps/fred-agents is the reference agent pod. Its templates are available
to any team after enrollment via the control plane. Third-party pods can supply
additional templates using the same mechanism.
| Template ID | Name | RAG | Description |
fred.github.assistant |
Custom assistant |
Optional |
Build-your-own ReAct agent; operator chooses the tools (document search, data analysis, ...) and writes the prompt at enrollment — no default capabilities pre-selected |
fred.github.rag_expert |
RAG expert |
Yes |
ReAct with Knowledge Flow vector search; formats cited sources with [N] anchors |
fred.github.deep_assistant |
Custom planning assistant |
Optional |
Plans before it acts on the LangGraph deep-agent runtime — breaks complex requests into steps and checks intermediate results; operator chooses tools and prompt; no filesystem tools by default |
fred.github.react_rag_mcp |
ReAct RAG (MCP) |
Yes |
Document-grounded ReAct agent backed by Knowledge Flow MCP search (vs. rag_expert's built-in tool ref); Tools-tab configurable library, search policy, and RAG scope |
fred.github.sql_expert |
SQL expert |
No |
Natural-language to SQL; table exploration via MCP database tool |
fred.github.sentinel |
Sentinel |
No |
Observability and monitoring assistant; MCP servers locked — operator-configured only |
fred.github.test_assistant |
Test harness |
N/A |
No-LLM deterministic agent covering all SSE event types, HITL flows, markdown rendering, error scenarios; used for integration and UI testing |
fred.github.platform_ops |
Platform ops |
No |
ReAct admin agent defaulting to the platform_postgres capability (see Agent capabilities); raised per-turn tool-call cap (30, vs. 12 default) for unattended multi-query investigation, with guardrails steering it to ground on the discovered schema, aggregate in SQL, and fix rather than blindly retry failed queries |
ReAct execution engine
All production agents use a LangGraph-backed ReAct loop. The graph is stateless
across requests; all state lives in the LangGraph checkpoint.
| Property | Detail |
| Framework | LangGraph agent graph |
| Cycle | think → plan → tool_call → observe → reflect → finalize |
| Thought kinds |
planning · tool_use · observation · reflection · synthesis — each visible as a status event and in the reasoning trace UI |
| Model routing | Configurable model profile per agent instance; tunable via prompts.system override field |
| KPI per node | Latency, token counts, and tool timings tracked per graph step |
| Langfuse tracing | Full trace exported with user_id, team_id, agent_instance_id, session_id, checkpoint_id, template_agent_id |
MCP tool integration
Agents use tools via the Model Context Protocol. The MCP catalog is declared in the
agent pod's mcp_catalog.yaml and proxied to the frontend by the control
plane — the control plane does not interpret tool behavior, only routes configuration.
| Capability | Detail |
| Catalog declaration | Per-pod mcp_catalog.yaml — server IDs, config fields, behavioral contracts, optional agent_instructions fragment injected into system prompt |
| Per-server config | Typed config_fields[] with FieldSpec (string/number/boolean/enum, required, default, validation); set per agent instance at enrollment |
| Server selection | Agent instance holds tri-state: inherit pod defaults · empty list · explicit subset of server IDs |
| Locked servers | Servers marked locked=True cannot be toggled by team managers — operator-controlled only |
| Bundled servers | Knowledge Flow text search, Knowledge Flow corpus; extensible via custom MCP server deployment |
Agent capabilities
Beyond MCP-backed tools, agent pods ship native capabilities registered via
fred.capabilities entry points. Each declares a team_scope:
ADMIN_GATED (default) requires a platform admin to enable it per team;
DEFAULT_ON capabilities are usable by every team out of the box.
| Capability ID | Tool(s) | Description |
html_artifact |
render_html_artifact(title, html, css) |
Renders agent-authored static HTML/CSS in a read-only side panel (Preview/HTML/CSS tabs, download). Capped at 256 KB combined HTML+CSS. ReAct-only. Rendered client-side in a sandbox="" <iframe srcDoc> (no allow-scripts, no allow-same-origin), with author markup passed through DOMPurify and an injected default-src 'none' CSP — no script execution and no network egress on any output path. |
platform_postgres |
postgres_list_tables(), postgres_run_query(sql) |
Read-only SQL access to the platform's own database, governed rather than raw. Read-only is enforced server-side in three layers — single-statement parsing (rejects stacked statements before execution), an explicit READ ONLY transaction, and a connect-time read-only session default — not by client-side SQL validation. Results are capped at 200 rows and 1000 chars/cell; query timeout is configurable, 1-120s. Bundled by the platform_ops agent template (see Bundled agent templates). |
document_similarity |
find_similar_passages(anchor, document_uids, top_k?) |
Ranks passages most similar to a given anchor passage within one or more explicitly named documents — a targeted comparison primitive, not corpus-wide search; at least one target document is required. Calls Knowledge Flow's vector-search similarity endpoint: embedding-based retrieval over a widened candidate pool, reordered by a cross-encoder. Configurable default_top_k (1-50), rerank, min_score. |
document_verbatim |
read_document(uid, ...) |
Positional, literal document read, one bounded page at a time ("what does the first paragraph say?", "read section 3") — distinct from a lossy summary. |
document_extract |
extract_from_document(uid, criterion) |
Exhaustive enumeration by criterion ("list ALL the requirements"), run server-side in Knowledge Flow as a map-reduce over every chunk rather than paged into the agent's own context. Gated behind a human proceed/cancel confirmation by default (require_confirmation, per-instance configurable) — the gate runs before extraction starts, so a cancel spends no extraction tokens. |
team_wiki |
wiki_list_pages(), wiki_read_page(path), wiki_propose_page(title, content_md, parent_path?), wiki_propose_page_text(path, content_md), wiki_publish_proposal(id) |
Read and propose changes to the team's wiki. Every write is a human-approved proposal, never direct — wiki_publish_proposal only submits it for approval, and is HITL-gated. ReAct-only: the rules page reaches the model through a prompt fragment a Graph agent's own composer does not support. |
Markdown rendering
Chat messages are rendered with react-markdown +
remark-gfm + rehype-sanitize. A streaming fence
guard prevents transient parse errors during chunked arrival.
| Feature | Detail |
| CommonMark base | Bold, italic, headings, lists, blockquotes, horizontal rules, links |
| GFM extensions | Tables, strikethrough, task lists, autolinks |
| Code blocks | Syntax highlighting; fenced blocks with language tag |
| Math | KaTeX inline ($...$) and block ($$...$$) |
| Diagrams | Mermaid (` ```mermaid `) rendered as SVG |
| Directives | :::details collapsible blocks; mindmap JSON blocks |
| Citations | [N] markers converted to clickable SourceBadge atoms linked to source detail modal |
| Images | Presigned MinIO URLs (1-minute TTL) injected server-side; rendered inline |
| Security | rehype-sanitize applied; no raw HTML injection |
Reasoning trace
Every agent turn renders an expandable reasoning trace alongside the answer, showing
the full thought process — plans, tool calls with arguments and results, observations,
and reflections.
| Behaviour | Detail |
| Auto-open | Trace panel opens on the first status event during streaming |
| Auto-close | Panel collapses on the final event |
| Entry types | Combo: tool_call + matching tool_result grouped as one card. Solo: plan, thought, observation, error as individual cards. |
| Status chips | Pending (streaming) · ok (success) · error — shown on each tool card |
| Detail drawer | Monaco JSON viewer showing full {call, result} or solo ChatMessage payload |
| Sources panel | Sources extracted from tool_result and final events; expandable with document viewer link when source.uid present |
Reasoning selector
A composer control showing which model is actually routed for the next
turn, plus a boolean toggle for whether that model reasons — not a
low/medium/high effort picker.
| Property | Detail |
| Wire field | RuntimeContext.reasoning: bool | None — tri-state: unset (agent doesn't offer reasoning at all), false (offered, left off), true (requested). A permission, not a guarantee: a platform-level ceiling can still suppress it. |
| Reasoning effort | Not user-controlled. Fixed per model by the ops-authored thinking-profile config and only displayed, never chosen per turn — a per-turn effort override was tried and withdrawn because some providers reject unsupported effort values. |
| Model identity | The chip displays the model actually routed for the next turn, independent of whether reasoning is offered on it. |
| Component | ReasoningChip in the composer's bottom row (widget id reasoning_toggle), sourced from ExecutionPreparation.chat_controls[] — only rendered for widgets the active agent actually exposes. |
| Persistence | Session-scoped only (stored client-side per session id), default off; not a global or account-level preference. |
Chat attachments
Files attached in the chat composer are fast-ingested into Knowledge Flow for
session-scoped retrieval, then persisted as session attachment records via the
control plane so they survive across turns and reloads.
| Feature | Detail |
| Fast ingest | POST /knowledge-flow/v1/fast/ingest — multipart file + session_id + scope: "session" + options_json; returns document_uid and a preview summary_md |
| Persistence | POST /teams/{team_id}/sessions/{session_id}/attachments records the attachment (name, mime, size, summary, document_uid) against the session; listed via the matching GET, removed via DELETE .../attachments/{attachment_id} |
| Context injection | Persisted attachments are listed by name (and internal document_uid) in a Markdown block prepended to the agent's system context, so document search tools can find them — no raw file path is passed |
| Multimodal | Images ≤ 4 MB of type png/jpeg/webp/gif are additionally read client-side and carried inline (data URL) as vision context for capable models |
| Progress UI | Local Redux task events (taskRegistered / taskEventReceived) drive per-attachment ingesting/ready/error state — no task SSE stream or polling involved for attachments |
Sessions
| Property | Detail |
| Ownership split | Control plane owns: title, timestamps, status, agent/team binding. Runtime owns: message content and checkpoints. |
| Session list | GET /teams/{team_id}/sessions — ordered by updated_at DESC, limit 50 |
| Editable titles | SessionTitleEditor component; title defaults to first user message (up to 120 chars) |
| Deep links | /team/{teamId}/managed-chat/{agentInstanceId}?session={uuid} |
| Purge policies | Control-plane-managed session retention; session_purge_queue handles scheduled deletion |
Prompt library & marketplace
Every team has a private, team-scoped prompt library. On top of it, a
cross-team marketplace lets teams publish, discover, use, and import
each other's prompt templates.
| Endpoint | Description |
POST /teams/{team_id}/prompts/{prompt_id}/publish | Sets a live published flag on the team's own prompt row — the marketplace listing reads that same row, so edits and the usage counter propagate immediately, no review step. Team-editor only; personal-space prompts cannot be published (400). |
POST /teams/{team_id}/prompts/{prompt_id}/unpublish | Withdraws it from the marketplace; the team keeps its own private copy unchanged. |
GET /marketplace/prompts | Lists every published prompt platform-wide, most-used first. Not team-scoped — open to any authenticated user. |
GET /marketplace/prompts/{prompt_id} | Full text of one published prompt, fetched on demand when its card is opened (listing payload carries only a preview). |
POST /marketplace/prompts/{prompt_id}/use | Records a "use" (copy to clipboard) by incrementing the prompt's shared usage counter. |
POST /marketplace/prompts/{prompt_id}/import | Copy-by-value into one or more of the caller's own teams: a fresh prompt row (new id, version 1, counters reset, unpublished) per target; name collisions get an _imported-N suffix. |
Team wiki Beta
Every team has a private, structured knowledge base — pages nested in a tree, each
with a versioned history. Humans write and edit directly; agents can read the whole
tree, but every agent-authored change is a proposal, never a direct write, until a
human explicitly approves it.
Beta. Shipped to gather real usage before a broader rollout —
expect it to be enabled for a subset of teams or deployments rather than
everywhere at once, and for details on this page to still move.
| Property | Detail |
| Structure | Pages nested up to a fixed depth, each addressed by a stable slug independent of its title. A dedicated "rules" page — read by every agent with wiki access before it answers — cannot have children. |
| Versioning | Append-only revisions: an edit never overwrites, it publishes a new revision and moves the page's current pointer. Full history with restore, paginated with a keyset cursor rather than an offset. |
| Conflict handling | Every write is a compare-and-swap against the revision its author actually read. One that moved on since is refused with the current text to rebase onto, never silently overwritten. |
| Agent authoring model | An agent reads with wiki_read_page and proposes with wiki_propose_page / wiki_propose_page_text (see Agent capabilities) — no tool writes directly. A published agent proposal leaves the page flagged for review until a human clears it. |
| Proposal lifecycle | A human approves (wiki_publish_proposal, HITL-gated) or declines a proposal. One nobody acts on expires automatically after a configurable retention window (30 days by default), so it cannot be published later by an unrelated retry. |
| Access | team_editor writes and approves; any team member reads. Same team-scoped ReBAC boundary as the rest of the platform (see Team scoping). |
Role model
Fred enforces two orthogonal role axes. Platform roles are global, held as
direct ReBAC relations on the platform-wide organization
singleton (see Team scoping). Team roles are
per-relationship (OpenFGA tuples) on each team. Neither axis grants access
to the other's surfaces — this is enforced at the API layer, not only in
the UI.
| Role | Axis | Authority |
owner |
Team (ReBAC) |
Create teams, assign managers, set TeamPlatformPolicy (quotas, allowed models, allowed MCP servers) |
manager |
Team (ReBAC) |
Configure TeamRoutingPolicy, manage team agents (enroll/edit/delete instances), manage team-scoped prompts |
member |
Team (ReBAC) |
Use team agents, use visible prompts, manage own personal-team prompts and agents |
platform_admin |
Platform (ReBAC) |
Full org-level administration — access internal/admin APIs, grant/revoke platform_observer. Only the bootstrap-root admin may grant/revoke platform_admin itself, and the root cannot be revoked. |
platform_observer |
Platform (ReBAC) |
Read-only platform-level access; grantable and revocable by any platform_admin |
Orthogonality rule: owner cannot manage agents or prompts.
manager cannot touch platform policy or quotas. Hardcoded in the authorization layer.
Platform-role management
| Endpoint | Description |
GET /users/platform-roles | List every platform-role holder, plus whether the caller is the bootstrap root |
POST /users/{user_id}/platform-roles | Grant platform_admin or platform_observer to a user |
DELETE /users/{user_id}/platform-roles/{relation} | Revoke a platform role from a user |
Imputability audit logs: every grant and revoke emits a structured audit-log
event (authz.relation.granted / authz.relation.revoked) carrying the
acting user, the subject, the relation, and the resource — a durable record of who assigned
what to whom.
Team scoping
| Concept | Detail |
| Personal team |
Every user has exactly one personal team. ID derived deterministically from Keycloak UUID via personal_team_id(user_id). Cannot be shared; governed by the same authorization model as regular teams. |
| Team-scoped resources |
Agents, prompts, sessions, document libraries, storage quotas — all scoped to a team. There is no global/unscoped resource namespace. |
| Organization singleton |
organization:fred in OpenFGA holds global role context without granting implicit team access. Used for platform-admin checks. |
| FrontendBootstrap |
GET /control-plane/v1/frontend/bootstrap returns resolved current_user, active_team, available_teams, permissions, feature_flags, ui_settings, gcu_version — single authenticated round-trip for the SPA shell. |
| PermissionSummary |
Flattened boolean capability flags per team (can_manage_team_agents, can_create_prompt, etc.) — UI renders conditionally on these, not on role strings. |
Policies & quotas
| Policy | Set by | What it controls |
TeamPlatformPolicy |
Owner |
Maximum storage quota, allowed model profiles, allowed MCP servers, team-level rate limits |
TeamRoutingPolicy |
Manager |
Default model profile for new agent instances, default RAG scope, search policy |
| Storage quota |
Owner (default: 5 GB team, 5 GB personal) |
Enforced on knowledge-flow document uploads per team |
| MCP server locks |
Pod author / platform |
Servers marked locked=True cannot be toggled by managers; operator-configured only |
| GCU (usage terms) |
Platform |
Version-tracked acceptance state per user; returned in FrontendBootstrap.gcu_version |
Runtime API
| Endpoint | Description |
POST /agents/execute/stream | SSE streaming execution (primary path) |
POST /agents/execute | Synchronous execution |
GET /agents/sessions/{session_id}/messages | Full turn history as ChatMessage[] |
GET /agents | Agent template catalog (pod-scoped) |
GET /health | Liveness / readiness probe |
Control-plane API
| Endpoint | Description |
GET /control-plane/v1/frontend/bootstrap | Single-call SPA init — user, team, permissions, flags |
POST /teams/{team_id}/agent-instances/{id}/prepare-execution | Issue runtime execute/stream URLs + effective chat options — no signed grant, the pod authorizes each request itself |
GET/POST/PATCH/DELETE /teams/{team_id}/agent-instances | Enroll, list, update, remove agent instances |
GET/POST/PATCH/DELETE /teams/{team_id}/sessions | Session metadata CRUD |
GET /teams/{team_id}/agent-templates | Proxied catalog from all registered runtime pods |
POST/GET /api/v1/tasks | Start and list long-running tasks (ingestion, migration) |
GET /api/v1/tasks/{id}/events | SSE task event stream with Last-Event-ID replay |
POST /api/v1/tasks/{id}/cancel | Cancel task (idempotent, 202) |
OpenAI compatibility layer
A secondary interface that allows tools expecting the OpenAI Chat Completions API
to connect to Fred agents without modification.
| Endpoint | Notes |
GET /v1/models | Model list (mapped to registered agent templates) |
POST /v1/chat/completions | Streaming-compatible; X-Fred-Team-Id header for team scoping |
Limitation: team-scoped managed execution and HITL are not fully supported
via the /v1 protocol. For full platform features use the native runtime API.
fred-sdk
The shared contract library imported by agent pods and the runtime. It contains no
server or product dependency — any team can build against it to author a deployable
agent pod.
| Module | Provides |
| Execution types | ExecutionTarget, RuntimeExecuteRequest, ChatMessage, VectorSearchHit, HumanInputRequest, all SSE event types |
| UiPart primitives | LinkPart, GeoPart and the UiPart discriminated union |
| ThoughtKind enum | planning · tool_use · observation · reflection · synthesis |
| Prompt utilities | System prompt assembly, tuning field resolution, context injection helpers |
| Authoring primitives | Base classes and decorators for defining agent behaviors in custom pods |
Observability
| System | What it captures | Configuration |
| Langfuse |
Full LLM traces with identity context: user_id, team_id, agent_instance_id, session_id, checkpoint_id, template_agent_id |
Enabled via LANGFUSE_* env vars; traces per turn, per node |
| Prometheus |
Process metrics, SQL connection pool metrics, per-agent KPI counters, latencies |
observability.metrics: prometheus in deployment config |
| KPI pipeline |
In-process counters (token counts, latency per step, tool timings) exported at configurable intervals |
kpi_process_metrics_interval_sec, kpi_log_summary_interval_sec |
| Correlation IDs |
request_id, trace_id, correlation_id propagated through all log lines, KPI entries, and metrics |
Automatic; no configuration required |
| Structured logs |
JSON-structured log output at configurable level (debug default); all services |
app.log_level in deployment config |
Security & compliance
Authentication
| Feature | Detail |
| Identity provider | Keycloak — OpenID Connect, configurable realm |
| Auth flows | PKCE browser flow (primary), resource-owner password flow (CLI/dev), no-security mode (local/airgapped) |
| Token validation | Bearer token (Keycloak JWT) required on all protected endpoints, including execution — authorized per-request via OpenFGA, no separate grant token |
| MFA | TOTP secrets migrate with Keycloak realm export; WebAuthn/passkey supported (hardware-bound, re-enrolment required on migration) |
| No-security mode | Auth bypass for local development and air-gapped deployments; team_id defaults to "personal" |
Authorization
| Feature | Detail |
| Engine | OpenFGA — Relationship-Based Access Control (ReBAC) |
| Policy evaluation | Runtime checks tuple relationships on every protected operation; no role string comparison in business logic |
| Tuple format | user:<keycloak-uuid> format exclusively (no username strings in production) |
| Enforcement layer | API level — not UI-only guards. Even direct API calls are rejected if tuples do not authorize the relationship. |
Data isolation
| Control | Detail |
| Team namespace isolation | All data (agents, prompts, sessions, documents) is team-scoped at the storage and authorization layer |
| Storage scope | Object storage paths are scoped by team_id at request time — an OpenFGA check by the pod handling the request, not a field in a signed grant |
| Session content isolation | Runtime returns session messages only to requests whose Keycloak JWT passes the per-request OpenFGA check for that session's team |
| Personal team isolation | Personal team data is accessible only to the owning user; the personal team ID is not guessable (UUID-derived) |
| Document download TTL | Presigned URLs for document images expire after 1 minute |
Security certification
C3 homologation achieved. Fred has completed its first homologation
at the C3 level of French national security classification. Security posture is
actively maintained: architecture reviews, OpenFGA policy audits, and Keycloak
hardening are ongoing as the platform evolves toward new homologation cycles.
Design for regulated environments
| Property | How it is achieved |
| Auditability | All execution attributed to user_id + team_id; Langfuse traces record full identity context per turn |
| Role separation | Platform (platform_admin/platform_observer) and team (team_admin/team_editor/team_analyst/team_member) roles are orthogonal and API-enforced; no privilege escalation path |
| Least-privilege execution | Each execution request is authorized independently by the pod via a per-request OpenFGA check against the caller's Keycloak JWT — scoped to the caller's team, cannot be forged or escalated by the frontend, and there's no grant object to extend or replay |
| No shared global state | Every resource is team-scoped; there is no unscoped global API for non-admin operations |
| Air-gap capable | No-security mode + local object storage + local vector store; no mandatory external service dependencies in the execution path |
Deployment
Kubernetes native
| Concern | Approach |
| Service exposure | Standard Kubernetes Service (internal) + Ingress/Gateway (browser); runtime pods exposed at /runtime/{runtime_id} prefix |
| Service discovery | Kubernetes DNS; no custom mesh or sidecar required |
| Network policy | Standard NetworkPolicy primitives; each service is independently policy-able |
| Helm chart | Modernized chart per FRED-CHART-MODERNIZATION-RFC.md; values-driven deployment for all services |
| Custom agent pods | Import fred-runtime, package as container, register runtime URL in platform.runtime_catalog_sources; no Fred core changes needed |
Configuration model
| Layer | File | Contains |
| Secrets / env | ENV_FILE | Database passwords, API keys, Keycloak secrets |
| Deployment config | CONFIG_FILE (YAML) | Base URL, ingress prefix, runtime catalog sources, processing profiles, log level, metrics config |
| Processing profiles | configuration.yaml → processing.profiles | PDF engine, OCR settings, chunk size, input processor registry per suffix |
| Audio config | configuration.yaml → audio_model | Whisper model size, device (cpu/cuda), language override |
Local development
| Command | What it does |
make run | Start the service's API process |
make run-worker | Start the Temporal/background worker |
make cli | Backend validation CLI (all backends expose this) |
make code-quality | Ruff + format (Python) or tsc + prettier (frontend) — run from repo root |
make test | Offline unit tests only (no live stack required) |
Infrastructure stack (local): PostgreSQL, Keycloak, MinIO (object storage),
OpenSearch (vector index), OpenFGA (authorization), Temporal (background workers) —
all provided via Docker Compose from ignored/fred-deployment-factory.