Self-Hosting and Operations
Self-Hosting and Operations is an instructor-led VDF AI course for IT and platform engineers who run VDF AI packages inside their own infrastructure. In four live half-days you mint portal credentials, plan compute, storage, network, identity and secrets, bring up and validate a package and set its operating routine, then earn the VDF AI Certified Platform Engineer certificate.
- 4 live half-days
- 6 modules + capstone
- Remote or on-site
- VDF AI Certified Platform Engineer
- Level
- Advanced
- Format
- Live and instructor-led, remote or on-site
- Length
- Four live half-day sessions (3.5 hours each)
- Audience
- IT operations, platform and SRE engineers, and the security staff who approve self-hosted deployments
- Cost
- Free for customers and partners; quoted for other teams
- Certificate
- VDF AI Certified Platform Engineer
- Reply to applications
- Within 2 business days
- Labs
- One after every module
What you will be able to do
- Take a portal account from registration to active and mint pull credentials safely
- Size compute, storage and the database for a pilot and for production
- Plan ingress, TLS, SSO and secrets around the infrastructure you already run
- Bring up a package and validate it against five explicit checks
- Deploy into an air-gapped environment through your internal registry
- Connect logs, health checks and metrics to your observability stack and set an update rhythm
Prerequisites
- Working knowledge of Linux, containers and either Docker Compose or Kubernetes
- An active portal.vdf.ai account and a host or cluster you are allowed to deploy to
6 modules and a capstone
The modules run across the four sessions. Each one ends with a lab in VDF AI.
-
The portal, packages and credentials
How your organisation gets access to self-hosted packages, and why the portal never reaches into your environment.
- Account states: pending verification, pending approval, active, trial, suspended and expired
- The package catalogue: versions, channels, system requirements, services and install guides
- Time-limited, read-only pull credentials and the limit on how often they are minted
- Sharing login and pull commands with the platform team for a deploy window
- One-way trust: the portal does not deploy, monitor or receive telemetry
- Lab
- Read a package’s system requirements and service list in the catalogue, generate credentials, authenticate your container runtime and pull the package images before the credentials expire.
- Outcome
- Pull access that is short-lived, read-only and recorded against the person who minted it.
-
Sizing, platforms and orchestration
Choosing the hosts, runtime and orchestrator that fit a pilot now and production later.
- Single-host baseline sizing compared with production sizing across several nodes
- Tested Linux distributions, multi-architecture images and why desktop platforms stay out of production
- Container runtimes: Docker, Podman and containerd
- Docker Compose, Kubernetes manifests or a Helm chart, and other OCI orchestrators
- PostgreSQL 16 or later: bundled for pilots, managed or clustered for production
- Lab
- Write sizing and orchestration plans for a pilot and for production in your environment, including the database tier and the usage evidence that would justify scaling up.
- Outcome
- A deployment plan that starts small and states what would trigger the move to production sizing.
-
Network, TLS and identity
Fitting a package into your existing network edge and identity provider rather than building new ones.
- Outbound TLS to the package registry, with no inbound ports opened on your side
- Restrictive network policies between the services inside a package
- Ingress through your own load balancer or reverse proxy, with certificates from your own authority
- SSO at the reverse proxy with SAML or OIDC, alongside per-instance user accounts
- The air-gapped path: a connected staging host and your internal registry
- Lab
- Design the ingress, certificate and SSO path for your deployment and the image route into an internal registry for an air-gapped site, then review both against the pre-deploy checklist.
- Outcome
- An edge and identity design that keeps your identity provider and certificates under your control.
-
Storage, secrets and observability
The foundations that decide whether a deployment can be run, watched and restored.
- Container volumes, database storage and SSD-backed disks
- Snapshot backups of the database and persistent volumes, and a restore drill
- Secrets in an environment file for pilots or in your secrets manager for production
- JSON logs on standard output and error, per-service health endpoints and Prometheus-compatible metrics
- Lab
- Map the package’s secrets into your secrets manager, point your log collector and metrics scraper at its services, and plan a daily database snapshot with a restore drill.
- Outcome
- Secrets, storage and telemetry that fit the tooling your team already operates.
-
Bringing up and validating a package
From pulled images to a healthy deployment you can hand over to users.
- The recommended startup configuration: compose file, environment file and install guide
- A first bring-up with Docker Compose on a single host
- Five checks: healthy services, a reachable frontend, a first admin, clean logs and a working basic flow
- Moving to Kubernetes once the first deployment is validated
- Hand-off: user accounts, connected sources and integrations
- Lab
- Bring up a package on a single host, work through all five validation checks, create the first admin user and note any repeating errors you find in the logs.
- Outcome
- A validated deployment with recorded evidence for each check.
-
Operating a self-hosted deployment
Running a deployment week after week: updates, credentials, access and the operating cadence.
- Pulling updates on a deliberate cadence or following an update channel
- The catalogue page as the place for current versions and release notes
- Minting fresh credentials for new hosts and new cluster nodes
- Log retention, backup schedule and monitoring as the operating cadence
- Trial expiry and account states, and what keeps running when the portal is unreachable
- Lab
- Write the runbook for your deployment, covering the update cadence, credential routine, log retention, backup schedule and monitoring alerts.
- Outcome
- A runbook your platform team can follow without the person who did the first deployment.
Capstone: a production plan for your own environment
Produce a deployment design for your organisation covering sizing, orchestration, database, ingress, TLS, SSO, secrets, observability and updates, bring a package up in a lab environment and walk a VDF AI engineer through your validation evidence and runbook.
VDF AI Certified Platform Engineer
Awarded to IT and platform engineers who complete Self-Hosting and Operations and pass the capstone review.
- Managing portal access and short-lived pull credentials
- Planning compute, storage, network, identity and secrets
- Bringing up and validating a self-hosted package
- Operating updates, observability and backups
What your team gets
- Four live half-day sessions (3.5 hours each)
- A hands-on lab after every module
- A materials pack: session slides and lab guides
- A capstone review with a VDF AI engineer
- The VDF AI Certified Platform Engineer certificate on passing the capstone
Apply, agree dates, learn
- Send the application below. It takes two minutes.
- We reply within 2 business days. Then we agree dates.
- Your team gets the materials pack. Then the live sessions begin.
Related courses: Platform Administration and Governance; VDF AI API and Integration Engineering.