platformengineering/openchoreo platformengineering conference wso2

Core Idea

The SREday talk’s argument: decouple the platform into planes and pluggable modules, abstract Kubernetes in three tiers, and expose the whole thing to AI through Model Context Protocol.

  • The platform engineering dilemma and how OpenChoreo enters as the answer.
  • Architecture: decoupled planes, pluggable modules, and the three tiers of Kubernetes abstraction (developer, platform, architecture via cell-based).
  • AI-native platforming: MCP integration, the developer experience shift, and live demonstrations.

Internal Developer Platforms (IDPs), Platform Engineering, Kubernetes Abstractions, and Model Context Protocol (MCP) Integration

Speaker: Lakmal Warusawithana (Maintainer of OpenChoreo)
Event: SREday Seattle 2026 Q2
[!video]- OpenChoreo: Building AI-Native, K8s-First Platforms | Lakmal Warusawithana | SREday Seattle 2026 Q2

1. Introduction & The Platform Engineering Dilemma

The Problem Space

In the modern software development landscape, developers face severe operational friction that limits their core productivity:

  • The Context-Switching Tax: Research demonstrates that developers spend more than 90% of their day-to-day time context-switching between disparate tools (e.g., build systems, CI/CD pipelines, security scanning tools, and observability suites).
  • CNCF Landscape Fatigue: When organizations turn to Platform Engineering to solve this friction, they encounter massive complexity. Within any single functional category in the Cloud Native Computing Foundation (CNCF) landscape, there are typically 20 to 30 projects to evaluate, learn, integrate, and maintain.
  • The "Do-It-Yourself" (DIY) [[Internal Developer Platform (IDPs)]] Failure:
    • Building an Internal Developer Platform (IDPs)) from scratch using tools like Backstage typically requires a dedicated platform engineering team up to two years of engineering effort.
    • Because the cloud-native ecosystem evolves so rapidly, DIY platforms are often outdated by the time they are delivered.
  • The AI Gap: Generative AI tools allow developers to write application code exponentially faster. However, moving this code into production remains bottlenecked by human platform engineering workflows.

Enter OpenChoreo

OpenChoreo is an open-source, β€œbatteries-included” platform engineering framework designed to resolve the DIY integration bottleneck.

  • Origin: Developed as a commercial SaaS tool (β€œChoreo”) and matured over five years with a large enterprise user base.
  • Governance: Donated directly to the Cloud Native Computing Foundation (CNCF) as an official Sandbox Project. The project recently celebrated its v1.2 release.
  • Value Proposition: Empowers platform engineers to deliver self-service infrastructure portals to developers on Day One, providing high-level abstractions without forcing teams to learn underlying Kubernetes YAML.

2. Decoupled Architecture: Planes & Modules

OpenChoreo is constructed using a highly decoupled architecture split into specialized functional Planes and pluggable Modules. This design prevents vendor lock-in and allows multi-cloud operations.

OpenChoreo Overview > Architecture

The Architectural Planes

  • Control Plane: The centralized management interface. It exposes standard UI views, Command Line Interfaces (CLIs), and Model Context Protocol (MCP) servers to developers and other organization personas.
  • Data Plane: The execution boundary where workloads run. It abstractly defines logical environments (e.g., Dev, Staging, Prod) which can be mapped directly to single clusters or mapped multi-cluster across multi-cloud and multi-region infrastructure.
  • Observability Plane: Responsible for aggregating and indexing distributed logs, metrics, and telemetry traces.
  • CI (Build) & CD (Deployment) Planes: Responsible for handling source-to-image code compilation, security scanning, and automated application delivery pipelines.

Pluggable Modules (Abstractions over Tooling)

OpenChoreo does not bind the platform to specific underlying technologies. Instead, it utilizes interfaces called Modules:

  • Example: The CD Plane uses an abstraction interface. Underneath that interface, platform engineers can swap FluxCD for alternative GitOps engines.
  • Example: The Observability Plane defines telemetry interfaces that can easily swap underlying technologies (e.g., open-source Prometheus/Grafana vs. commercial APM platforms) without changing the developer-facing abstraction layer.
  • Current Community Progress: The community is currently developing 18 core project modules, 12 of which are fully stabilized and completed.

3. The Three Tiers of Kubernetes Abstraction

To bridge the operational gap between developers (who want to write code) and platform engineers (who manage infrastructure), OpenChoreo implements three progressive abstraction layers:

Tier 1: Developer Abstractions

Developers do not write Kubernetes manifests or configure ingress rules. They interact solely with high-level business concepts:

  • Project: A logical grouping of functional applications.
  • Component: A single operational application block (e.g., an API microservice, a database, or a static frontend web UI).
  • Endpoint: An interface exposed by a component to receive incoming traffic.
  • Dependency: Resources that a component consumes (e.g., an external database or a message broker).
  • Configuration: Environment variables and runtime configuration files associated with the workload.

1. What is Platform Engineering > Introduction

Tier 2: Platform Abstractions

Platform engineers maintain full access to the granular power of Kubernetes. They map developer concepts into underlying cloud-native resources:

  • Platform engineers define Component Types matching enterprise standards.
  • They write structural templates mapping simple developer declarations back to custom resource definitions, limits, and runtime primitives.

Tier 3: Architecture Abstractions (Cell-Based Architecture)

OpenChoreo dynamically maps design-time architecture (like Domain-Driven Design) to concrete runtimes using a paradigm called Cell-Based Architecture (CBA).

    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€ CELL BOUNDARY ──────────────────────────┐
    β”‚                                                                    β”‚
    β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”             β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”               β”‚
    β”‚  β”‚  Component A    │◄───────────►│  Component B    β”‚               β”‚
    β”‚  β”‚  (Microservice) β”‚             β”‚   (Database)    β”‚               β”‚
    β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜             β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜               β”‚
    β”‚           β–²                                                        β”‚
    β”‚           β”‚ (Internal Traffic Allowed)                             β”‚
    β”‚           β–Ό                                                        β”‚
    β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”                                               β”‚
    β”‚  β”‚   API Gateway   │◄──────────────────────────────────────────────┼(External HTTPS)
    β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ (Explicit perimeter control / Zero Trust)     β”‚
    β”‚                                                                    β”‚
    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
  • Definition of a Cell: A Cell represents a runtime-enforced bounded context mapping to a software subdomain.
  • Implementation: Under the hood, OpenChoreo configures a Cell using dedicated Kubernetes namespaces paired with automated Network Policies.
  • Zero-Trust Network Perimeter:
    • All internal components within a single Cell can communicate freely.
    • Any traffic aiming to cross the boundary of a Cell must go through an explicitly declared API Gateway component.
    • This guarantees that traffic entering or exiting a bounded context is authenticated and authorized at runtime, building zero-trust perimeter security without manual network engineering.

4. AI-Native Platforming & Model Context Protocol (MCP)

OpenChoreo shifts away from traditional, click-heavy user portals to adopt an AI-Native interface model.

Integrating the Model Context Protocol (MCP)

OpenChoreo exposes two core MCP Servers:

  1. Control Plane MCP Server: Allows LLM agents to orchestrate platform-level operations (creating resources, environments, pipelines, and managing RBAC).
  2. Observability Plane MCP Server: Allows LLM agents to query telemetry, search logs, fetch active component metrics, and review distributed trace spans.

The Developer Experience Shift

  • IDE-Centric Orchestration: Developers no longer need to switch back and forth from their terminal/IDE to a web-based console. They interact with their preferred coding agents (e.g., Cursor, VS Code assistants) using natural language.
  • Built-in Specialized Agents: OpenChoreo features an internal family of autonomous micro-agents:
    • SRE Agent: Analyzes alert anomalies, logs, and traces.
    • FinOps Agent: Monitors and recommends optimizations for cloud expenditures.
    • Architecture Agent: Validates structural deployments against design-time DDD blueprints.
    • Remediation Agent: Generates and executes configuration patches to resolve system failures.

5. Walkthrough of Live Demonstrations

Demo 1: Dynamic Site Maps and Traffic Overlays

  • The Portal UI: Built using a customized Backstage portal connected to the OpenChoreo control plane.
  • Dynamic Visualization: The UI automatically renders a comprehensive visual map of the infrastructure topology (clusters, data planes, active pipelines, environments, and services).
  • Overlaying Design vs. Runtime:
    • By combining design-time declarative data with real-time telemetry tracing, the platform overlays live traffic vectors directly onto the architectural model.
    • Anomaly Detection:
      • Orphan Services: Highlights services running in the cluster that are drawing compute resources but are receiving zero real-world traffic.
      • Security Violations (The β€œDoor in a Red” Indicator): Instantly flags live runtime traffic paths between components that have no design-time architectural declarations. This warns engineers of a potential system compromise or structural rule violation.

Demo 2: Prompt-Driven Platform Provisioning

  • The Task: A platform engineer is tasked with onboarding two brand-new software teams:

    • Payment Service Team: Requires Dev, Integration, and Production logical environments.
    • Customer Portal Team: Requires Staging, UAT, and Production pipelines.
  • The Manual Friction: In standard enterprises, provisioning these environments requires writing hundreds of lines of YAML, managing namespace permissions, setting up separate pipeline definitions, and updating IAM policiesβ€”a process that takes days.

  • The OpenChoreo Solution:

    • The engineer provides a single natural language prompt detailing the multi-environment, multi-pipeline onboarding requirements to their MCP-enabled agent.
    • The agent translates the prompt, coordinates with the OpenChoreo Control Plane APIs, spins up the logical namespaces, designs the corresponding CD pipelines, and binds the environments to the respective teams in under a minute.

Demo 3: Automated Root Cause Analysis (RCA) & Remediation

  • The Induced Failure: The speaker introduced a database connection string misconfiguration inside an API service (corrupting a PostgreSQL endpoint URL). The application frontend broke immediately.

  • Step 1: Automated Detection & Incident Creation

    • The platform’s observability loop instantly caught the cascading HTTP 500 errors and created an incident inside the dashboard.
  • Step 2: Automated Triage (RCA Agent)

    • An SRE sub-agent activated. Instead of requiring manual log searches, the agent queried the Observability Plane MCP.
    • The agent evaluated the broken state against a historical ledger of configuration diffs, system changes, and DNS errors.
    • Within 30 seconds, the agent produced a clear Root Cause Analysis report with high confidence, identifying the exact malformed database connection string.
  • Step 3: Human-In-The-Loop Remediation

    • The Remediation Agent analyzed the configuration diff and generated a precise JSON configuration patch to restore the correct connection string.
    • Security Guardrail: The remediation agent operates under strict Role-Based Access Control (RBAC) permissions (e.g., only authorized to modify environment variables and basic configuration maps).
    • The platform prompts a human engineer to verify the patch inside the portal. Once the engineer clicks β€œApply,” the system deploys the configuration change and the application automatically recovers.

6. Key Takeaways & Architectural Lessons

  • AI Pragmatism: AI should not be applied blindly. Standard promotions (e.g., promoting code from Dev to Staging) are best left to a single-button click or automated GitOps webhooks. However, AI is incredibly effective for multi-layered platform actions (like provisioning complex environments) and high-stress scenarios (like rapid triage of 3:00 AM production outages).
  • Contextual Feed is Vital: For an AI agent to achieve high-confidence operational troubleshooting, it must be fed rich local platform context, including system dependencies, historical configuration changes, and recent deployment diffs.
  • Clean Data Isolation: Dividing OpenChoreo into separate logical data planes enables platform groups to manage centralized control policies while dynamically dividing underlying hardware resources to meet localized latency, compliance, or cost demands.