How do you integrate on-prem AI with Entra ID, Okta, or Keycloak?
Use SAML or OIDC for authentication, map identity-provider groups to platform roles, and enforce permissions across agents, tools, models, retrieval sources, admin actions, and environments. The audit trail should record the authenticated actor and the policy decision for each request. In VDF AI, Entra ID sign-in and group-to-role mapping are built in; Okta and Keycloak connect through an SSO-aware reverse proxy.
What happens in a full air-gapped AI environment?
The platform runs without outbound internet. Models, containers, patches, documentation, and evaluation artifacts are imported through an inspected offline process. Inference, retrieval, embeddings, logs, and admin actions stay inside the disconnected environment.
How does observability integrate with our SIEM?
The AI platform should export structured events for prompts, retrieval, tool calls, model routes, approvals, responses, failures, and policy decisions. Security teams can retain, search, alert, and investigate those events in their existing SIEM.
Is private RAG the same as enterprise search?
No. Enterprise search finds documents. Private RAG retrieves permitted context, passes it to the model, captures citations, enforces source permissions, manages embeddings and indexes, and records the evidence used to produce an answer.
Can on-prem AI use cloud models at all?
Yes, if policy allows it. A hybrid architecture can route low-risk workloads to approved external models while pinning sensitive workloads to local models. The important requirement is that egress is explicit, policy-gated, and logged.
What are the best on-prem GPU servers for running self-hosted agentic AI platforms?
Match the server class to the model tier the router will actually use: L4/RTX-class nodes for small routed models, a single L40S (48 GB) server for 7B–14B pilot workloads, A100/H100 80 GB servers for 70B-class production serving with vLLM, and 4×–8× H100/H200 nodes when many teams or 100B+ open-weight models consolidate onto one platform. Buy for the sustained utilization floor and let routing absorb peaks — an idle GPU fleet is how the on-prem cost case dies.
How do I size CPU, GPU, and VM capacity for an LLM agent platform?
Use the two-formula starting point, then benchmark: GPU memory ≈ model parameters × bits per weight ÷ 8 plus KV-cache (≈ 2 × layers × KV heads × head dim × active tokens × bytes), and CPU/VM capacity ≈ 2–4 cores with 8–16 GB RAM per concurrent agent session for orchestration, retrieval and policy checks. Then replace the estimate with a benchmark of the full governed workflow — identity, retrieval, reranking, inference, tool calls, and log export — because agent platforms bottleneck on the workflow, not on raw tokens per second.
How should we size GPU capacity for on-prem LLM inference?
Start from representative workflows and measure end-to-end latency, not only model tokens per second. Include concurrency, context length, output length, retrieval fanout, reranking, tool latency, batch jobs, and failover targets before buying production hardware.
What audit fields should a governed on-prem LLM platform keep?
At minimum: actor, role, timestamp, request purpose, policy decisions, data sources retrieved, model and version, prompt or redacted prompt, tool calls, approval actions, response, evaluation signals, errors, and retention metadata.
How is this different from Azure OpenAI private endpoints or AWS Bedrock with private networking?
Private cloud networking can reduce exposure to the public internet, but the model service remains cloud-managed. A true on-prem or air-gapped architecture keeps the production data plane, model runtime, retrieval indexes, logs, and governance evidence under customer-controlled infrastructure.