msflib-ai-core¶
Purpose¶
msflib-ai-core is the AI infrastructure layer for MSFLib host apps. It gives you provider-agnostic factories for chat models, embeddings and vector stores, a prompt registry with scoped overrides, an ingestion (chunking) pipeline, usage and rate-limit helpers, and LangGraph checkpoint and memory policies. ai-api and knowledge build on it, and conversation calls it optionally (memory cleanup).
It does not run a model server and does not decide your product's prompts or retrieval logic. It resolves which provider, model, prompt and store apply to the caller's scope, and builds them.
For a hands-on walkthrough, see AI capabilities.
Install¶
[tool.poetry.dependencies]
msflib = { git = "https://github.com/msflib/fastapi.git", subdirectory = "core", rev = "core-v0.2.1" }
msflib-ai-core = { git = "https://github.com/msflib/fastapi.git", subdirectory = "modules/ai_core", rev = "ai_core-v0.2.2", extras = ["openai", "pgvector"] }
Provider and backend packages are optional extras, so a bare install stays light. Selecting a provider or backend whose package is missing raises an ImportError (ModuleNotFoundError). Most backends name the extra to install; pgvector and the chat providers report only the missing package (for example langchain_postgres), so use the table below to find the extra.
| Extra | Enables |
|---|---|
openai, azure |
langchain-openai, for the provider ids openai and azure_openai (azure alone is not a provider id) |
anthropic, ollama, bedrock, groq, openrouter |
The matching chat provider |
huggingface |
langchain-huggingface; hosted Inference API embeddings use the provider id huggingface_api |
pgvector |
pgvector vector store (langchain-postgres, psycopg) |
qdrant, pinecone, chroma, weaviate, milvus, faiss |
That vector store backend (client plus langchain integration; faiss uses faiss-cpu and langchain-community) |
checkpoint-postgres |
Postgres LangGraph checkpointer and memory store |
llamaindex, community |
Extra chunking and document-loading providers |
mcp |
MCP-backed tool loading |
ddgs or tavily |
A search backend for the web_search tool (pick one, per WEB_SEARCH_BACKEND) |
web |
httpx, required by the fetch_url tool |
The authoritative list is [tool.poetry.extras] in modules/ai_core/pyproject.toml. Each vector backend has an extra named after it.
What is in the package¶
| Area | Package path | What it does |
|---|---|---|
| Settings | msflib.ai_core.config |
AICoreSettings, the AI_CORE settings namespace |
| Model factories | msflib.ai_core.providers |
get_llm, get_embeddings, the provider type registry, register_provider_type |
| Vector stores | msflib.ai_core.vector_store |
get_vector_store, scoped retrieval helpers, backend registry |
| Ingestion | msflib.ai_core.services.ingestion_pipeline, .services.chunking, .services.extraction |
Load, clean and chunk documents; pluggable chunking and loader registries |
| Prompts | msflib.ai_core.services.prompt_registry |
Versioned prompts with user, workspace and global overrides |
| Provider and store profiles | msflib.ai_core.models, .services.provider_registry, .services.vector_store_registry |
Database-backed, scope-resolved provider and vector-store configs, with encrypted keys and optional failover |
| Runtime policy | msflib.ai_core.services.config_bridge |
Turns resolved settings into LangChain middleware and agent policy |
| LangGraph | msflib.ai_core.services.checkpoint_policy, .memory_store_policy |
Scope-aware checkpointer and long-term memory store |
| Tools | msflib.ai_core.tools |
current_datetime_tool, calculator_tool, web_search_tool, fetch_url_tool, rag_search_tool and tool registries |
| Tracking | msflib.ai_core.tracking |
TokenUsageCallback, RateLimiter |
| HTTP API | msflib.ai_core.router |
Provider-profile and vector-store-profile routers (member and admin) |
Note that msflib/ai_core/__init__.py exports nothing: import from the subpackages shown above.
Wiring into a host app¶
Settings are read from the AI_CORE namespace of your settings object. get_ai_dependencies returns plain callables for programmatic use, and get_provider_dependencies / get_vector_store_dependencies return request-scoped FastAPI dependencies for routers.
from msflib.ai_core.deps import get_ai_dependencies
from msflib.ai_core.config import AICoreSettings
from msflib.core.config import SettingsBase
class HostSettings(SettingsBase):
AI_CORE: AICoreSettings = AICoreSettings(LLM_API_KEY="sk-...", EMBEDDING_API_KEY="sk-...")
ai = get_ai_dependencies(HostSettings())
llm = ai.get_llm()
embeddings = ai.get_embeddings()
In your own app, pass the SettingsBase instance you already have; it only needs to carry the AI_CORE scope. This is not the standalone AICoreSettings instance used in the AI capabilities guide. The *_from_registry variants additionally take a session and a scope (see Scopes, tenancy and workspaces).
Known issue (#291)
Importing msflib.ai_core.models (which the registry services and routers import) registers aiproviderprofile and vectorstoreprofile, whose tenant_id references tenant. If msflib.tenancy.models.tenant has not been imported, create_all and migrate fail with NoReferencedTableError. Import msflib.tenancy.models.tenant before creating tables. ai_prompt_version has no tenant reference, and the package root msflib.ai_core registers nothing by itself.
The namespace members are get_llm, get_embeddings, get_vector_store, their *_from_registry variants (which resolve database profiles for the caller's scope), get_middleware_config and get_summarization_config.
To expose the profile-management HTTP API, mount the routers from msflib.ai_core.router. Each is a factory that returns an APIRouter and takes keyword arguments only. get_session is a callable FastAPI dependency (not a Session object), and settings is your host settings object. The identity callables differ by router:
| Factory | Required identity callables | Optional | Default prefix |
|---|---|---|---|
provider_router |
get_current_workspace, get_current_workspace_anonymous, get_current_active_user, workspace_role_check |
get_current_tenant, get_policy_resolver, prefix, tags |
/ai/providers |
vector_store_router |
get_current_workspace, get_current_active_user, workspace_role_check |
get_current_tenant, get_policy_resolver, prefix, tags |
/ai/vector-stores |
provider_admin_router |
get_current_active_superuser |
get_current_tenant, prefix, tags |
/ai/admin/providers |
vector_store_admin_router |
get_current_active_superuser |
get_current_tenant, prefix, tags |
/ai/admin/vector-stores |
See Mounting routes for the general pattern, and the dependency table under "Endpoint dependency wiring" in modules/ai_core/README.md for which identity callables to pass in your host.
Configuration¶
Set values through the AI_CORE namespace. How environment variables map depends on how your host settings class includes the module:
- If the host declares the settings as a field (
AI_CORE: AICoreSettings = AICoreSettings()), use the double-underscore form for any setting:AI_CORE__LLM_PROVIDER,AI_CORE__EMBEDDING_MODEL,AI_CORE__RATE_LIMIT_RPM. - If the host mixes
AICoreSettingsinto its settings class (the style the testsite uses), use the flat setting name:LLM_PROVIDER,EMBEDDING_MODEL,RATE_LIMIT_RPM. The double-underscore form is ignored in this style. - In either style, eight settings also accept an
AICORE_-prefixed alias:AICORE_LLM_PROVIDER,AICORE_LLM_MODEL,AICORE_LLM_API_KEY,AICORE_EMBEDDING_PROVIDER,AICORE_EMBEDDING_MODEL,AICORE_VECTOR_STORE_BACKEND,AICORE_MIDDLEWARE_POLICYandAICORE_CHUNKING_POLICY. No other setting has an alias; use the form from the first two bullets for everything else.
Behaviour may change (#280)
The NAMESPACE__KEY style only works for composed (field) settings, not for subclassed hosts.
| Key | Default | Notes |
|---|---|---|
LLM_PROVIDER, LLM_MODEL |
openai, gpt-4o |
LLM_PROVIDER is validated against the provider type registry; LLM_MODEL is a free string |
LLM_API_KEY, LLM_BASE_URL |
empty, None |
|
LLM_TEMPERATURE, LLM_MAX_TOKENS |
0.0, None |
Temperature must be 0.0 to 2.0 |
EMBEDDING_PROVIDER, EMBEDDING_MODEL |
openai, text-embedding-3-small |
|
EMBEDDING_API_KEY, EMBEDDING_BASE_URL |
empty, None |
|
VECTOR_STORE_BACKEND |
pgvector |
|
VECTOR_STORE_URL, VECTOR_STORE_API_KEY |
None |
VECTOR_STORE_URL is required by pgvector, qdrant, weaviate and milvus; pinecone needs VECTOR_STORE_API_KEY instead; chroma and faiss need neither |
VECTOR_STORE_COLLECTION_PREFIX |
msflib |
Naming prefix for collections, not a tenant boundary |
VECTOR_STORE_EXTRA |
{} |
Backend-specific options (FAISS path, Pinecone index and region, and so on) |
RETRIEVER_COLLECTION_NAME, RETRIEVER_K |
documents, 4 |
RETRIEVER_K must be 1 to 50 |
RATE_LIMIT_RPM |
60 |
Must be at least 1 |
TOKEN_BUDGET_PER_WORKSPACE |
None |
Read by the provider registry but not enforced by RateLimiter or anything else (#278) |
DB_PROVIDER_REGISTRY_ENABLED, PROVIDER_FAILOVER_ENABLED |
True, True |
|
PROVIDER_SECRET_ENCRYPTION_KEY |
None |
Encrypts API keys stored in provider profiles. When unset it falls back to the core SECRET_KEY; set it explicitly in production so rotating one does not break the other. SECRET_KEY is random per process by default, so if neither is set, stored provider keys become unreadable after a restart |
DB_VECTOR_STORE_REGISTRY_ENABLED, VECTOR_STORE_FAILOVER_ENABLED |
True, False |
Vector failover is opt-in because it changes which documents a search can see |
WEB_SEARCH_BACKEND, WEB_SEARCH_API_KEY |
ddgs, empty |
tavily needs a key |
Some keys can be overridden per request or per scope (the allowlist is request_override_allowlist in config.py). The full set of fields, including the web-search, task-profile and local-disk-backend options, is in AICoreSettings.
Extension points¶
register_provider_type(ProviderTypeSpec(...), aliases=...)adds a chat or embedding provider that langchain'sinit_chat_modeldoes not cover.register_vector_store_backend(...)adds a vector store backend.ChunkingRegistry.register_strategy(...)adds a chunking strategy; seemodules/ai_core/msflib/ai_core/services/chunking/README.md.UnifiedLoaderRegistryadds document loaders; seemodules/ai_core/msflib/ai_core/services/extraction/README.md.- The
ToolRegistryprotocol lets you supply your own tool sets, including MCP-backed ones.
Troubleshooting¶
| Symptom | Cause and fix |
|---|---|
ImportError or ModuleNotFoundError |
The provider or backend package is not installed. Find the matching extra in the table above. |
ImportError: cannot import name 'get_llm' from 'msflib.ai_core' |
Import from the subpackage: from msflib.ai_core.providers import get_llm. |
Unsupported LLM_PROVIDER: ... (a validation error from settings) or ValueError: Unsupported LLM provider: ... (from get_llm) |
The name is not in the provider type registry. Use a registered provider id (for example azure_openai, not azure) or call register_provider_type. |
| Pgvector errors about a missing type | Run CREATE EXTENSION vector; on the database once. |
Checkpointer or memory store backend="postgres" fails |
Install the checkpoint-postgres extra, pass connection_string=, and call with run_setup=True once at startup (it defaults to False). |
| Search returns documents from another workspace | You called similarity_search directly. Use similarity_search_scoped so tenant and workspace filters are mandatory. |
API reference¶
See the generated API reference for msflib.ai_core. modules/ai_core/README.md remains the detailed reference for provider profiles, runtime policy, LangGraph scoping and the tools library.