Skip to content

msflib-ai-api

Purpose

msflib-ai-api is the HTTP transport for AI features in a host app. One router gives you scoped retrieval (/search), retrieval-augmented answers (/ask), and a tool-using agent (/agent/ask, with a streaming variant), all resolving the caller's account, workspace and tenant, and delegating model, vector-store and tool behaviour to ai-core.

It is deliberately thin. It has no settings of its own (it reads AI_CORE), no upload endpoint (files go through documents and are indexed there), and no persistence beyond what the agent's checkpointer and the optional transcript writer provide.

Install

[tool.poetry.dependencies]
msflib = { git = "https://github.com/msflib/fastapi.git", subdirectory = "core", rev = "core-v0.2.1" }
msflib-ai-api = { git = "https://github.com/msflib/fastapi.git", subdirectory = "modules/ai_api", rev = "ai_api-v0.2.1" }

This pulls in msflib-account, msflib-ai-core and msflib-tenancy. Model providers, vector backends and optional tools come from ai-core's extras (for example openai, pgvector, mcp); see ai-core. The package msflib/ai_api/__init__.py is empty, so import from the submodules shown below.

Wiring into a host app

Add AICoreSettings and AccountSettings to your settings class, then mount the router. get_current_account is required, because ai-api has no anonymous surface, and must return an account-like object with id, role and current_workspace_id.

from msflib.ai_api.router import router as create_ai_router

app.include_router(
    create_ai_router(
        settings=settings,
        get_session=get_session,
        get_current_account=get_current_account,
        get_current_workspace=get_current_workspace,   # optional
        get_current_tenant=get_current_tenant,         # optional
        prefix="/ai",
    ),
    prefix="/api/v1",
)

The routes, relative to the router prefix:

Route Purpose
POST /search Scoped similarity search over the vector store.
POST /ask Retrieval-augmented answer.
GET /status Module, workspace, configured provider and model, vector backend.
GET /protocols Response protocols /ask supports.
GET /agent/status Agent mode, source (builder, injected or injected_with_tools), provider and model.
POST /agent/ask Run the agent once and return its output.
POST /agent/ask/stream Same request, returned as a text/event-stream.

router() composes two factories, create_llm_router (search, ask, status, protocols) and create_agent_router (the /agent routes), both importable from msflib.ai_api.router if you want only one. The prefixes llm_prefix (default empty) and agent_prefix (default /agent) place them under the main prefix.

Without get_current_tenant, the tenant with slug default is looked up in the database and must already be seeded, otherwise requests fail with 500. Without get_current_workspace, requests run with workspace_id=None, the workspace-less scope.

Known issue (#288)

The default-tenant fallback ignores the host's TENANCY.DEFAULT_TENANT_SLUG. If you changed the slug, requests return 500 "Default tenant has not been seeded" even though your seeding used the same settings. Pass get_current_tenant built from get_tenant_dependencies(session_dep=get_session, settings=settings.scope("TENANCY")).get_current_tenant.

See Mounting routes for the general pattern.

Router parameters worth knowing

Parameter Effect
model_name_or_factory Pin the model. A string such as "openai:gpt-4o" is resolved once at mount time. A zero-argument callable is called lazily on first use and cached. With llm_factory_per_request=True it is called on every request with session=, scope= and ai_settings= keyword arguments (a string is rejected then).
get_vector_store Supply the vector store (anything with similarity_search(query, *, k, filter)). Without it, the store comes from AI_CORE settings.
tiered TieredResolution(vector_store=..., web_search_backend=...). Opt in to re-resolving those settings per request through the tenant, workspace and user policy tiers, at the cost of a database lookup per request. Both default to False.
get_policy_resolver The policy resolver used for tiered settings.
get_injected_agent Supply your own agent. A zero-argument callable is used as is. A callable taking an argument receives the composed tool list. It must return a LangChain Runnable.
get_tool_registries, get_mcp_servers, mcp_load_timeout Control the agent's tools. See below.
get_knowledge_engine Adds knowledge graph tools to the default tool set.
system_prompt A string, or a zero-argument callable, for the built agent.
checkpoint_backend, checkpoint_connection_string LangGraph checkpointer for the built agent. Defaults to "memory".
memory_backend, memory_connection_string LangGraph long-term memory store. Defaults to "memory".
transcript_writer_factory, speaker_label_factory Persist turns into a conversation and label speakers.

Configuration

ai-api has no settings namespace. Model, embedding, vector-store, retrieval and web-search behaviour come from AI_CORE; see the table in ai-core. The environment variable names depend on how you compose the settings:

  • Subclassed (class Settings(AICoreSettings, ...)): the flat names LLM_PROVIDER, LLM_MODEL, VECTOR_STORE_BACKEND, RETRIEVER_K and WEB_SEARCH_BACKEND, (a few also accept AICORE_ aliases such as AICORE_LLM_PROVIDER). AI_CORE__LLM_PROVIDER is ignored in this style.
  • Nested field (AI_CORE: AICoreSettings = AICoreSettings()): AI_CORE__LLM_PROVIDER, AI_CORE__LLM_MODEL, AI_CORE__VECTOR_STORE_BACKEND, AI_CORE__RETRIEVER_K, AI_CORE__WEB_SEARCH_BACKEND. The flat RETRIEVER_K is not read.

One environment variable is read directly:

Variable Default Effect
SCOPE_CONSTRAINTS_ENFORCE unset 1 or true sets constraint enforcement to enforce, so a failed scope check returns 403. 0 or false forces log-only. Unset uses each operation's default, which is log-only.
SCOPE_DECISION_LOGGING_ENABLED unset 0 or false silences the per-request scope decision log.

Behaviour may change (#270)

Scope constraints default to log-only: a failed scope check is logged but not blocked unless you set SCOPE_CONSTRAINTS_ENFORCE. The default may change.

Known issue (#292)

With SCOPE_CONSTRAINTS_ENFORCE=1, every request that has a workspace is rejected with 403 "Access denied: workspace_not_member" when the workspace is resolved from the path or a header. The membership check reads the account's current_workspace_id, which a path- or header-resolved workspace does not update, so it is None. Until this is fixed, leave enforcement off (the default log-only mode logs MembershipConstraint ... workspace_not_member for each such request and lets it through). Do not copy the requested workspace id into current_workspace_id yourself: the check treats that field as proof of membership, so it would let an account satisfy the constraint for a workspace it does not belong to.

The answer model for /ask is chosen in this order: a model injected through model_name_or_factory; otherwise ai-core's provider registry, using the task name aiapi.ask, so a provider profile or AI_CORE.TASK_PROFILES binding for aiapi.ask (or a broader prefix such as aiapi, or default) routes answers to that provider; otherwise the static AI_CORE.LLM_* settings. If no model is available at all, /ask degrades to returning the top retrieved passage instead of failing.

Key concepts

Scope per request. For /search, /ask and the agent routes, ai-api builds a ScopeEnvelope from the request: tenant, workspace, the account (as the principal, with its role) and the X-AI-Channel header (default api), plus conversation_id and sub_thread_id from the body. Retrieval filters on that scope, so a caller only ever searches its own tenant and workspace. The scope is then checked against the aiapi.retrieval.search or aiapi.ask profile (tenant match, workspace membership, role and channel). The default level is log-only: violations are logged, not blocked, until you set SCOPE_CONSTRAINTS_ENFORCE (see the notice under Configuration). Workspace membership is judged from the account's current_workspace_id, and a workspace resolved from the path or a header does not update it, so keep it in step with the workspace dependency you pass before turning on SCOPE_CONSTRAINTS_ENFORCE (see the known issue under Configuration).

Request bodies. SearchRequest has query, k (1 to 50, default 4), document_type and source_id. AskRequest has question, k, document_type and source_id. Both accept optional conversation_id and sub_thread_id. A conversation id ties an agent turn to a checkpoint thread, so the next turn resumes it; without one, each agent call is single-turn.

Response protocols. /ask and /agent/ask return ai-api's native shape by default. Select another with ?protocol=openai or the X-AI-Protocol header; the header wins. The supported values are native, openai, langchain, ag_ui, a2ui, mcp, acp and a2a. An unknown value returns 400. Adapters are registered in a ProtocolAdapterRegistry (msflib.ai_api.adapter_registry); GET /protocols lists what is active.

The agent. By default the router builds a LangChain agent with create_agent, using the model "<AI_CORE.LLM_PROVIDER>:<AI_CORE.LLM_MODEL>" (or your injected model), a checkpointer, a memory store and a tool set, and caches it for the router's lifetime. The runtime context passed on every call carries the request's scope (AgentRuntimeContext), which scope-aware tools such as rag_search and the knowledge tools read; it is never taken from a checkpoint or from model output. Agents you inject with get_injected_agent should set context_schema=AgentRuntimeContext if they use those tools.

Tools. Without get_tool_registries, the default tools are the current date and time, a calculator, web_search (when ddgs or tavily is installed), fetch_url (when httpx is installed) and rag_search (when the vector store can be built), each skipped quietly when its dependency is missing, plus knowledge tools when you pass get_knowledge_engine and msflib-knowledge is installed. Passing get_tool_registries replaces that default set with your own ToolRegistry objects. get_mcp_servers appends MCP-backed tools in either case.

Streaming. /agent/ask/stream returns Server-Sent Events. Token-level streaming happens only when the resolved agent is a compiled LangGraph graph, which is always true for the built agent. A plain LCEL runnable is streamed as LangChain yields it. For a multi-node injected graph, message streaming surfaces tokens from every chat-model call in the graph, so keep internal-only model calls (classification, planning) out of the path the endpoint invokes. See also SSE and WebSockets.

Transcripts and speaker labels. transcript_writer_factory(scope, payload) returns an object with write_user_turn and write_agent_turn; ai-api calls it best-effort, so a failing writer is logged and never breaks the answer. speaker_label_factory(scope, payload) returns a label that prefixes the user's message so an agent in a shared conversation can tell people apart. The conversation module provides ready-made implementations.

Examples

Search and ask, with an injected store and model

This runs as written. It stands in a fake vector store and model so no provider or database is needed; in a real app you pass a real get_session and omit get_vector_store and model_name_or_factory so ai-core resolves them from settings.

from types import SimpleNamespace

from fastapi import FastAPI
from fastapi.testclient import TestClient
from langchain_core.documents import Document
from langchain_core.runnables import RunnableLambda
from msflib.account.config import AccountSettings
from msflib.ai_api.router import router as create_ai_router
from msflib.ai_core.config import AICoreSettings
from msflib.core.config import CoreSettings, SettingsBase


class Settings(AICoreSettings, AccountSettings, CoreSettings, SettingsBase):
    pass


class FakeStore:
    def similarity_search(self, query, *, k, filter=None):
        return [
            Document(
                page_content="Pump P-101 feeds separator V-201.", metadata={"source_id": "doc-42"}
            )
        ]


app = FastAPI()
app.include_router(
    create_ai_router(
        settings=Settings(),
        get_session=lambda: None,
        get_current_account=lambda: SimpleNamespace(id=1, role="user", current_workspace_id=None),
        get_current_tenant=lambda: SimpleNamespace(id=1),
        get_vector_store=lambda: FakeStore(),
        model_name_or_factory=lambda: RunnableLambda(lambda _: "P-101 feeds V-201."),
        prefix="/ai",
    ),
    prefix="/api/v1",
)

client = TestClient(app)

hits = client.post("/api/v1/ai/search", json={"query": "pump", "k": 3}).json()
print(hits["count"], hits["hits"][0]["page_content"])   # 1 Pump P-101 feeds separator V-201.

native = client.post("/api/v1/ai/ask", json={"question": "What feeds V-201?"}).json()
print(native["answer"])                                  # P-101 feeds V-201.

openai = client.post(
    "/api/v1/ai/ask",
    json={"question": "What feeds V-201?"},
    headers={"X-AI-Protocol": "openai"},
).json()
print(openai["output"][0]["content"][0]["text"])        # P-101 feeds V-201.

Your own agent

Inject any LangChain Runnable. Here a trivial one stands in for a real agent; the router calls it with {"messages": [...], "k": ..., ...} and returns whatever it produces:

from langchain_core.runnables import RunnableLambda

echo_agent = RunnableLambda(lambda value: {"echo": value["messages"][0]["content"]})

router = create_ai_router(
    settings=settings,
    get_session=get_session,
    get_current_account=get_current_account,
    get_current_workspace=get_current_workspace,
    get_injected_agent=lambda: echo_agent,      # zero-arg: used as is, no tools or context
    prefix="/ai",
)

With a mounted app, GET /api/v1/ai/agent/status reports agent_source: "injected", and POST /api/v1/ai/agent/ask with {"question": "hi"} returns {"echo": "hi"}. This was run against a SQLite-backed app with a seeded default tenant. To receive the tools and the scope-bearing runtime context, write get_injected_agent=lambda tools: build_my_agent(tools) (one positional parameter) and build the agent with context_schema=AgentRuntimeContext from msflib.ai_api.agent_runtime.

Giving the agent knowledge tools and a transcript

from msflib.conversation.integrations import ai_bridge
from msflib.knowledge.deps import get_knowledge_dependencies

knowledge_deps = get_knowledge_dependencies(settings, session_factory=lambda: Session(engine))

create_ai_router(
    settings=settings,
    get_session=get_session,
    get_current_account=get_current_account,
    get_knowledge_engine=knowledge_deps.get_knowledge_engine,
    transcript_writer_factory=ai_bridge.build_transcript_writer_factory(
        lambda: Session(engine), agent_id="assistant"
    ),
    prefix="/ai",
)

Not run end to end, since it needs a working model and the knowledge and conversation tables; each name was checked against the source. See Knowledge and Conversation.

Troubleshooting

Symptom Cause and fix
500 "Default tenant has not been seeded" No tenant dependency was passed and the tenant with slug default is missing. Pass get_current_tenant or seed on startup (see the known issue under Wiring if you changed the slug).
401 "Authentication required" get_current_account returned None or an account without an id.
503 "AI service configuration unavailable" The vector store or agent could not be built or invoked: a missing provider extra, a bad AI_CORE setting, or a backend that is down. The details are in the server log.
400 "Unsupported X-AI-Protocol value" Use one of the values from GET /protocols.
403 "Access denied: ..." SCOPE_CONSTRAINTS_ENFORCE is on and the scope check failed, for example because the account's current_workspace_id does not match the workspace dependency (see the known issue under Configuration).
/ask returns "Top matching context says: ..." No model was resolved, or generation failed. Check provider settings or model_name_or_factory.
/ask says no indexed matches were found The search returned nothing in this caller's scope. Check that documents were ingested under the same tenant and workspace.
Knowledge tools raise "require scope on the agent's runtime context" A custom injected agent was built without context_schema=AgentRuntimeContext.
No streaming tokens until the end The agent is not a compiled LangGraph graph, or an internal node is not on the streamed path. See Streaming above.
Fewer default tools than expected A tool whose optional dependency is missing is skipped silently. Install ddgs or tavily, or httpx.

API reference

See the generated API reference for msflib.ai_api. modules/ai_api/README.md has curl examples for each route.

See also