Documentation

The optional dashboard

Everything Pahlevan learns is already reachable through kubectl get -o yaml, a Prometheus scrape, or a log line. In practice that means a learned profile is a YAML status block nobody reads until something is denied, and the first time anyone looks at it is during an incident.

The dashboard is the answer to that: a browser view of what each workload does - the process tree, the learned file, network and syscall surface, the flow from learning to enforcement, and what was denied and why - drawn as diagrams rather than another table of rows.

It is also the one Pahlevan component a browser talks to, which makes it the one an attacker would most like to find. The rest of this page is mostly about that.

It is optional, and optional means absent

The dashboard is off by default, and off means the objects do not exist:

  • It is not in install.yaml. The release manifest is generated by scripts/gen-install.sh from an explicit list of files, and deploy/dashboard is not on it. A cluster that applies a release asset installs the agent and the operator, exactly as before.
  • It is not in the default Helm values. dashboard.enabled is false, and a default helm template renders zero dashboard objects: no scaled-to-zero Deployment, no unused ServiceAccount. An idle ServiceAccount with a cluster-wide read binding is still a credential sitting in the cluster waiting to be borrowed.
  • It is a separate image and a separate Deployment. A cluster that never enables it never pulls those bytes, and the agent image does not grow an HTTP server and a bundle of static assets it has no use for.
  • Nothing in the agent or the operator imports the dashboard package, so its code does not run in either of them even when it is installed.

hack/install/dashboard_test.go asserts each of those, so "optional" stays true rather than staying written down.

Enabling it

Both paths need a TLS pair first. The server refuses to start without one unless --insecure is passed on purpose, because every request to it carries a Kubernetes bearer token.

Helm

kubectl create secret tls pahlevan-dashboard-tls \
  --cert=dashboard.crt --key=dashboard.key -n pahlevan-system

helm upgrade --install pahlevan charts/pahlevan-operator \
  --namespace pahlevan-system --create-namespace \
  --set dashboard.enabled=true \
  --set-string dashboard.networkPolicy.allowedNamespaceLabels."pahlevan\.io/dashboard-client"=true

Use --set-string for the namespace labels. --set x=true produces a YAML boolean, and the API server rejects a NetworkPolicy whose label value is not a string.

The values worth knowing:

Value Default What it does
dashboard.enabled false Renders nothing at all when false.
dashboard.replicaCount 2 The dashboard holds no state, so a lost replica costs a reconnect.
dashboard.image.repository ghcr.io/obsernetics/pahlevan-dashboard Its own image, separate from the agent's.
dashboard.rbac.create true The ServiceAccount and the grant below. There is no value that widens the grant.
dashboard.tls.secretName pahlevan-dashboard-tls Must contain tls.crt and tls.key.
dashboard.audiences [] Token audiences to require. See Authentication.
dashboard.networkPolicy.enabled true Default-deny around the dashboard.
dashboard.networkPolicy.allowedNamespaceLabels {} Namespaces allowed to reach port 8443. Empty means nobody.
dashboard.networkPolicy.allowedNodeCIDRs [] Node CIDRs for kubelet probes, if your CNI enforces policy on host traffic.
dashboard.networkPolicy.apiServerCIDR 0.0.0.0/0 Narrow this to your control plane.

Raw manifests

kubectl apply -k deploy/base                      # agent and operator, as usual
kubectl create secret tls pahlevan-dashboard-tls \
  --cert=dashboard.crt --key=dashboard.key -n pahlevan-system
kubectl apply -k deploy/dashboard                 # the dashboard, opt-in

deploy/dashboard/rbac-namespaced.yaml is deliberately not in that kustomization. It is the single-namespace alternative to the cluster-wide read binding; see the RBAC section.

The RBAC it needs, and why the list is short

The dashboard's ServiceAccount holds five grants and nothing else:

API group Resource Verbs Why
authentication.k8s.io tokenreviews create Turn the browser's bearer token into a username and groups.
authorization.k8s.io subjectaccessreviews create Ask, per read, whether that user may see this object.
policy.pahlevan.io pahlevanpolicies get, list, watch Draw the policies.
policy.pahlevan.io containerprofiles get, list, watch Draw the learned surface.
policy.pahlevan.io attacksurfaces get, list, watch Draw the risk view.

Creating a TokenReview or a SubjectAccessReview asks the API server a question. It returns an answer, changes nothing in the cluster, and grants the caller no access of its own.

The list is short because every entry on it is something a request-handling bug could reach. Specifically absent:

  • No pods. A pod list would include every container's environment, which is where people put credentials. The dashboard does not need it to draw a profile, so it does not have it. A viewer who wants a pod list uses kubectl with their own credentials.
  • No secrets, ever.
  • No write verbs, no /status subresources. The dashboard cannot create, update, patch, or delete anything. Changing a policy or an enforcement mode stays a kubectl operation under the operator's own credentials. The server enforces this a second time in code: it registers reads and refuses every other method before a handler runs.
  • No wildcards. verbs: ["*"] on three CRDs is a write primitive, and a wildcard is how a short grant stops being auditable.
  • No cluster-admin, and no binding to a built-in role. The grant is written out in deploy/dashboard/rbac.yaml rather than aliased to system:auth-delegator, so an operator can audit it by reading one file instead of trusting that a built-in still means what they remember.

The read grant is cluster-wide because the dashboard serves whatever namespaces the viewer is allowed to see. That is safe only because of the SubjectAccessReview per read described next. If you want it narrower anyway, apply deploy/dashboard/rbac-namespaced.yaml and delete the pahlevan-dashboard-read ClusterRoleBinding: the dashboard then reads one namespace, and a viewer whose own RBAC covers the whole cluster still sees only that one, because a SubjectAccessReview can narrow what the dashboard fetches and never widen it.

How authentication works

There is no login page, no user database, and no session secret. There is nothing to leak, because the dashboard stores no credential of its own beyond the ServiceAccount token every pod has.

  1. The browser presents a Kubernetes bearer token on the request.
  2. The dashboard calls TokenReview to ask the API server who that token belongs to. An invalid or expired token ends there.
  3. For every read, the dashboard calls SubjectAccessReview against that user's identity: may this user get this resource in this namespace? The object is fetched with the dashboard's own service account, but nothing reaches the browser unless the API server says the viewer could have fetched it themselves.

The consequence is the property worth having: two people opening the same URL see different things, and each sees exactly what kubectl would have shown them. There is no second permission model to keep in step with RBAC, which means there is no second permission model to get wrong.

Set dashboard.audiences once you are minting tokens for the dashboard. With it empty, any token the API server recognises is accepted, including a projected token some unrelated pod holds - which turns the dashboard into a relay for credentials it was never meant to see.

Getting a token to paste, for a quick look:

kubectl create token my-user --duration=10m

In a real deployment the token comes from whatever your ingress or proxy already does for OIDC, and the dashboard checks it the same way.

TLS

The dashboard serves TLS only. Plaintext requires --insecure, which exists for a unit test or a sidecar that terminates TLS in front of it, and which has to be set on purpose.

There is no self-signed fallback. A dashboard that generates its own certificate teaches its users to click through a browser warning, and a user trained to click through the warning cannot tell your certificate from a proxy's.

With cert-manager:

apiVersion: cert-manager.io/v1
kind: Certificate
metadata:
  name: pahlevan-dashboard-tls
  namespace: pahlevan-system
spec:
  secretName: pahlevan-dashboard-tls
  dnsNames:
  - pahlevan-dashboard.pahlevan-system.svc
  - pahlevan-dashboard.pahlevan-system.svc.cluster.local
  issuerRef:
    name: your-issuer
    kind: ClusterIssuer

Without cert-manager, supply the pair yourself:

kubectl create secret tls pahlevan-dashboard-tls \
  --cert=dashboard.crt --key=dashboard.key -n pahlevan-system

The secret is mounted read-only at /etc/pahlevan/tls with mode 0400. Add --tls-client-ca to require a client certificate as well, which narrows who may open a connection at all before any token is checked.

Exposing it safely

The project ships no public exposure, and that is deliberate. The Service is ClusterIP, the chart has no value that changes its type, and there is no Ingress manifest anywhere in this repository. A NodePort would publish the dashboard on every node's address and a LoadBalancer would ask your cloud provider for a public IP, either of which would turn --set dashboard.enabled=true into publishing a view of every workload's behaviour. That is not a decision one flag should be able to make.

Reaching it is your deliberate act, in ascending order of exposure:

Port-forward, which is enough for most people and exposes nothing:

kubectl port-forward -n pahlevan-system svc/pahlevan-dashboard 8443:443

Your own ingress. Use the ingress controller you already run, with the authentication you already put in front of your other internal tools. Then:

  • Label its namespace so the NetworkPolicy lets it through: kubectl label namespace ingress-nginx pahlevan.io/dashboard-client=true. Without that label nothing reaches the dashboard at all, which is the intended default.
  • Keep TLS end to end. The dashboard speaks HTTPS to the backend, so configure your controller for an HTTPS upstream rather than terminating at the edge and sending the token onward in plaintext.
  • Put the ingress on an internal load balancer, or behind your VPN. A view of what every workload in the cluster does, and what has been denied to it, is a reconnaissance report for anyone who reaches it and finds a bug in the token check. An unreachable port cannot be probed for that bug.

The shipped NetworkPolicy is default-deny in both directions, with holes for the labelled namespace, DNS to the cluster resolver, and the API server. Narrow apiServerCIDR to your control plane once you know its address: as shipped, a compromised dashboard could open a connection to anything on 443 or 6443, including something outside the cluster.

What it cannot do

  • It cannot write. No create, update, patch, or delete, in RBAC or in code.
  • It cannot read pods, secrets, or anything outside the three Pahlevan CRDs.
  • It cannot show a viewer more than that viewer's own RBAC allows.
  • It cannot load an external origin. The Content-Security-Policy forbids inline script and every third-party origin, and assets are embedded in the binary. A dashboard that pulls a charting library from a CDN has made every viewer's browser trust a third party, which is not a trade this project gets to make on a user's behalf.
  • It cannot touch the kernel. No eBPF, no host namespaces, no host mounts, no capabilities, a read-only root filesystem, and runAsNonRoot. The agent stays the only privileged component.

The threat model, and the division of responsibility between the project and you, is in SECURITY.md.

See also