AI, RAG & Vector Search
This hub explains how to build AI systems that can retrieve evidence, isolate tenant data, control model behaviour and measure answer quality—not merely call an LLM API.
When should Qdrant, PostgreSQL search or SQL be used?
How are groundedness, citations and tool use evaluated?
Define the boundary, then choose the technology.
Use RAG
Knowledge changes, evidence matters and the answer must cite private or domain-specific sources.
Use search or SQL
The user needs exact records, aggregates, filters or deterministic business facts.
Add an agent
The system must choose tools or execute a bounded workflow with authorization and audit evidence.
Continue with practical guides.
Qdrant multitenancy: isolate retrieval, scoring and storage
A production design for tenant-safe Qdrant retrieval: enforce scope in trusted code, separate sparse scoring populations and promote heavy tenants deliberately.
Read the guideAI security · Anthropic CVPAnthropic CVP: reduced cyber blocks need stronger architecture
Anthropic expanded CVP into three tiers. The real deployment work is identity, authorization, isolation, egress control, retention and revocation.
Read the guideAI · Decision infrastructureOpenAI Decisions API: build a decision system, not a magic threshold
OpenAI's Decisions API turns text and images into probabilities, choices and scores. The hard part is the policy around the answer.
Read the guideAI infrastructure · Vector searchMigrate embeddings in Qdrant without breaking search
A production migration plan for new embedding models: named vectors or blue-green collections, dual writes, backfill, shadow evaluation, cutover and rollback.
Read the guideAI infrastructure · Anthropic APIAnthropic Models API `line`: stop parsing model IDs
Anthropic added a model-family field to the Models API. Here is how to use it for discovery without confusing family, capability and release policy.
Read the guideAI · Content provenanceOpenAI textGrain: what text watermarking can—and cannot—prove
OpenAI made text watermarking opt-in for select API models. This guide explains textGrain, detection limits and a safe provenance architecture.
Read the guideAI · RAG chunkingRAG chunking for complex PDFs: preserve structure, tables and context
A production guide to layout-aware PDF chunking, table handling, parent-child retrieval, overlap, contextual metadata and measurable RAG evaluation.
Read the guideAI · Speech generationGemini 3.8 TTS in production: voices, streaming and consent
A production guide to Gemini 3.8 Flash TTS and Flash-Lite covering model routing, streaming audio, reusable voices, consent, caching and evaluation.
Read the guideAI agents · Runtime architectureGPT-6 agent loops: async tools, steering and safe recovery
A production guide to GPT-6 async tool calling and mid-turn steering with pending-work registries, WebSocket recovery, approvals, budgets and observability.
Read the guideAI infrastructure · API latencyOpenAI Ultrafast mode: engineer latency, not just token speed
A production guide to routing GPT-6 Astra Ultrafast requests with WebSockets, latency budgets, rate limits, residency checks, fallbacks and cost controls.
Read the guideAI · Model routingGPT-6.1 Sol in production: route by evidence, not hype
A practical architecture for adopting GPT-6.1 Sol with model routing, prompt caching, bounded tools, eval gates, fallbacks and cost-quality telemetry.
Read the guideAI developer tools · Claude CodeClaude Code Mods: event middleware for safer AI coding
Anthropic’s new Mods can rewrite Claude Code events and UI. Here is how to use them safely, test their order and choose between Mods, hooks, Skills and MCP.
Read the guide