ENGINEERING TOPIC HUB

AI, RAG & Vector Search

This hub explains how to build AI systems that can retrieve evidence, isolate tenant data, control model behaviour and measure answer quality—not merely call an LLM API.

QUESTIONS THIS HUB ANSWERS
01

How should documents be chunked and versioned?

02

When should Qdrant, PostgreSQL search or SQL be used?

03

How are groundedness, citations and tool use evaluated?

DECISION FRAMEWORK

Define the boundary, then choose the technology.

Use RAG

Knowledge changes, evidence matters and the answer must cite private or domain-specific sources.

Use search or SQL

The user needs exact records, aggregates, filters or deterministic business facts.

Add an agent

The system must choose tools or execute a bounded workflow with authorization and audit evidence.

CURATED GUIDES

Continue with practical guides.

All articles →
AI infrastructure · Vector search

Qdrant multitenancy: isolate retrieval, scoring and storage

A production design for tenant-safe Qdrant retrieval: enforce scope in trusted code, separate sparse scoring populations and promote heavy tenants deliberately.

Read the guide
AI security · Anthropic CVP

Anthropic CVP: reduced cyber blocks need stronger architecture

Anthropic expanded CVP into three tiers. The real deployment work is identity, authorization, isolation, egress control, retention and revocation.

Read the guide
AI · Decision infrastructure

OpenAI Decisions API: build a decision system, not a magic threshold

OpenAI's Decisions API turns text and images into probabilities, choices and scores. The hard part is the policy around the answer.

Read the guide
AI infrastructure · Vector search

Migrate embeddings in Qdrant without breaking search

A production migration plan for new embedding models: named vectors or blue-green collections, dual writes, backfill, shadow evaluation, cutover and rollback.

Read the guide
AI infrastructure · Anthropic API

Anthropic Models API `line`: stop parsing model IDs

Anthropic added a model-family field to the Models API. Here is how to use it for discovery without confusing family, capability and release policy.

Read the guide
AI · Content provenance

OpenAI textGrain: what text watermarking can—and cannot—prove

OpenAI made text watermarking opt-in for select API models. This guide explains textGrain, detection limits and a safe provenance architecture.

Read the guide
AI · RAG chunking

RAG chunking for complex PDFs: preserve structure, tables and context

A production guide to layout-aware PDF chunking, table handling, parent-child retrieval, overlap, contextual metadata and measurable RAG evaluation.

Read the guide
AI · Speech generation

Gemini 3.8 TTS in production: voices, streaming and consent

A production guide to Gemini 3.8 Flash TTS and Flash-Lite covering model routing, streaming audio, reusable voices, consent, caching and evaluation.

Read the guide
AI agents · Runtime architecture

GPT-6 agent loops: async tools, steering and safe recovery

A production guide to GPT-6 async tool calling and mid-turn steering with pending-work registries, WebSocket recovery, approvals, budgets and observability.

Read the guide
AI infrastructure · API latency

OpenAI Ultrafast mode: engineer latency, not just token speed

A production guide to routing GPT-6 Astra Ultrafast requests with WebSockets, latency budgets, rate limits, residency checks, fallbacks and cost controls.

Read the guide
AI · Model routing

GPT-6.1 Sol in production: route by evidence, not hype

A practical architecture for adopting GPT-6.1 Sol with model routing, prompt caching, bounded tools, eval gates, fallbacks and cost-quality telemetry.

Read the guide
AI developer tools · Claude Code

Claude Code Mods: event middleware for safer AI coding

Anthropic’s new Mods can rewrite Claude Code events and UI. Here is how to use them safely, test their order and choose between Mods, hooks, Skills and MCP.

Read the guide