Insights · Research · Field Notes

Enterprise AI, explained for builders and buyers. Page 2 of 27

Practical writing on governed agent orchestration, on-premise AI, compliance, and the infrastructure decisions that separate pilot projects from production platforms.

237 Articles
39 Topics
12 Featured
A long aisle between rows of metal server racks, the kind of self-hosted infrastructure where a RAG team keeps its vector index next to the documents it was built from
RAG 15 min read

Best Vector Database for RAG (2026): Open-Source and Self-Hosted Options

Ten vector databases checked against their own documentation and GitHub repositories in October 2026: pgvector, Qdrant, Milvus, Weaviate, Chroma, OpenSearch, Elasticsearch, LanceDB, Redis and Vespa, plus Pinecone as the managed reference. Compared on licence, hybrid search, filtering, quantization, multitenancy and the work it takes to run each one yourself.

#vector database#vector search#private RAG
Read article
A triple-fan graphics card standing upright on a white surface, the kind of single GPU that the smaller Gemma 4 sizes are built to run on
AI Infrastructure 11 min read

Gemma 4 On-Premise: Local Requirements, Sizes and Setup (2026)

Google's Gemma 4 family runs from phone-sized E2B to a 31B dense model, all under Apache 2.0. This guide lists every size and variant, the memory each needs for weights and KV cache, Google's official quantization-aware builds, and the install steps for vLLM, llama.cpp, Ollama, LM Studio and MLX.

#Gemma#open-weight models#local LLM
Read article
Out-of-focus lines of colour-highlighted code on a dark monitor, standing in for the terminal commands that install and serve gpt-oss on local hardware
AI Infrastructure 10 min read

gpt-oss Local Deployment Guide: Requirements and Setup (2026)

OpenAI's gpt-oss-20b and gpt-oss-120b are Apache 2.0 reasoning models that ship in MXFP4, so they fit one 16 GB or one 80 GB device. This guide covers the memory each size needs at real context lengths, the harmony format and reasoning effort, tool calling, and step-by-step setup with vLLM, Ollama and llama.cpp.

#gpt-oss#open-weight models#local LLM
Read article
A laptop on a wooden desk with its screen split into several tiled terminal and editor panes, next to a potted cactus and a coffee cup
AI Platforms 13 min read

Hermes Agent vs Claude Code vs OpenCode (2026): Agent Harnesses Compared

Hermes Agent, Claude Code and OpenCode compared from their own documentation, with OpenClaw alongside: licence, models and local inference, approval defaults, sandboxing, MCP, subscription rules and list prices, checked in October 2026.

#Hermes Agent#Claude Code#OpenCode
Read article
A hand gesturing over printed charts and a calculator on a desk, in front of a large monitor showing a chart, the kind of review where evaluation scores for a language-model release are compared
AI Infrastructure 12 min read

LLM Evaluation Tools (2026): Open-Source Frameworks Compared

promptfoo, DeepEval, Ragas, OpenAI Evals, lm-evaluation-harness, Inspect, MLflow, Langfuse, Arize Phoenix and TruLens compared from their repositories and docs in October 2026: licence, maintenance, offline and online evaluation, RAG and agent metrics, LLM-as-a-judge and CI integration.

#AI evaluation#open-source AI#AI governance
Read article
An engineer holding a laptop beside a row of server cabinets in a data centre corridor, representing a company running Mistral models on its own infrastructure
AI Infrastructure 11 min read

Mistral On-Premise: Open Models, Self-Deployment and Licensing (2026)

Mistral AI publishes some of its strongest models under Apache 2.0 and keeps others behind a revenue-capped or commercial licence. This guide sorts the October 2026 lineup by licence, explains Mistral's own self-deployment offers and data terms, sizes the GPUs each open model needs, and gives the vLLM, llama.cpp and Ollama steps to serve them.

#Mistral#open-weight models#local LLM
Read article
Black network cables plugged into the ports of a rack-mounted switch, a reminder that interconnect and power, not only GPU memory, separate a PCIe workstation card from an SXM data-centre GPU
AI Infrastructure 14 min read

RTX PRO 6000 vs H100 vs H200 vs B200 for LLM Inference (2026)

NVIDIA's RTX PRO 6000 Blackwell, H100, H200 and B200 compared from NVIDIA's own datasheets: memory, bandwidth, FP8 and FP4 Tensor Core throughput, NVLink or PCIe, power and form factor, plus sessions-per-GPU arithmetic and bandwidth-bound speed ceilings for 32B and 70B models.

#AI infrastructure#GPU capacity planning#LLM inference
Read article
Two colleagues in business clothes reviewing a printed document together in an office, representing a company drafting and approving its AI acceptable use policy
AI Governance 13 min read

AI Acceptable Use Policy Template: A Company AI Policy You Can Copy (2026)

A company AI policy you can copy and adapt: model wording for approved and prohibited tools, data classification, human review, customer disclosure under EU AI Act Article 50, AI literacy under Article 4, confidentiality, incident reporting, exceptions and review, with a rollout plan and a one-page checklist.

#AI governance#AI policy#EU AI Act
Read article
Five colleagues working on laptops at shared desks with plants and screen dividers in a bright open-plan office, the kind of staff an internal ChatGPT-style assistant is rolled out to
Implementation Guide 9 min read

How to Build an Internal ChatGPT for Your Company: Three Routes and a Checklist

Staff already use ChatGPT, often on personal accounts. This guide compares three ways to give them an internal ChatGPT instead: buy business seats, assemble an open-source stack, or deploy a platform on your own infrastructure. It ends with a build checklist.

#private AI#private RAG#local LLM
Read article
Stay ahead of enterprise AI

Turn insight into an on-prem AI roadmap.

Pair these articles with a product walkthrough to see how VDF AI handles orchestration, governance, and cost control inside your own infrastructure.