Releases

Changelog and packages

What changed in each Pahlevan release, and every artifact you can install: a distroless container image on GHCR, a Helm chart, and a single-file Kubernetes manifest. The current release is v3.6.0.

Packages

Three ways to install v3.6.0

One distroless image contains pahlevan-agent (the per-node DaemonSet), pahlevan-operator (the leader-elected Deployment), and the pahlevan CLI. No shell, no package manager, nothing but the binaries.

Helm chart

helm repo add pahlevan https://obsernetics.github.io/pahlevan/charts
helm repo update
helm install pahlevan pahlevan/pahlevan-operator \
  -n pahlevan-system --create-namespace

Raw manifest

kubectl apply -f https://github.com/obsernetics/pahlevan/releases/latest/download/install.yaml

The manifest carries the CRDs, RBAC, the operator Deployment, and the agent DaemonSet. Swap latest for v3.6.0 to pin a release.

Container image

docker pull ghcr.io/obsernetics/pahlevan:v3.6.0

Verifying the image

docker pull ghcr.io/obsernetics/pahlevan:v3.6.0
docker inspect ghcr.io/obsernetics/pahlevan:v3.6.0

# Digest to pin in air-gapped or regulated environments
docker inspect --format '{{index .RepoDigests 0}}' \
  ghcr.io/obsernetics/pahlevan:v3.6.0
Image tag Meaning
latest Most recent build of the default branch
v3.6.0 Immutable release tag, recommended for production
main Rolling tag for the default branch
main-<sha> Per-commit build of the default branch, useful for bisecting

Full package reference, including chart values and CRD cleanup, is in docs/packages.md.

Changelog

What changed, release by release

Following Keep a Changelog and semantic versioning. The canonical file is CHANGELOG.md.

3.6.0

2026-10-01 Current

Added

  • learningConfig.requireReview and learningConfig.reviewedAt. Learning is trust on first use: a workload already compromised when learning starts has its malicious behaviour baselined, and until now nothing required anyone to look at that baseline before it became the thing enforced. A policy that sets requireReview: true holds every matching container in learning once its window and grace period elapse, until an operator sets reviewedAt on the policy. Checked once, at the moment a container would otherwise transition: a baseline that changes after review is not re-reviewed. Off by default, so every existing policy keeps transitioning exactly as before.

Removed

  • pkg/cli.GetCodecs and pkg/grpcapi.EventTypeFromProto/EventTypeToProto: exported but called from nowhere in the tree, in a test, or in the docs. Each package already does its own type/event conversion inline where it is actually used.

3.5.1

2026-09-28

Fixed

  • ContainerProfile.Spec.PolicyRef is now populated. The field has existed since the type was added, but the live controller never set it when persisting a profile, so it was always empty. That silently broke profilesync's lookup of a policy's syscallPolicy overrides (they never reached a generated seccomp profile) and the CLI's per-policy grouping in pahlevan netpol generate and pahlevan attacksurface. A learning container now gets the field resolved live; an enforcing one gets it frozen at its enforce transition, matching the overrides and seccomp profile generated from that same decision.

Changed

  • Dependency batch: k8s.io/api, k8s.io/apimachinery, k8s.io/cli-runtime and k8s.io/client-go to 0.37.1, github.com/onsi/gomega to 1.44.0, and google.golang.org/grpc to 1.84.0. actions/download-artifact to v8 in CI.

3.5.0

2026-09-22

Releases are signed now, and the shipped manifests are sized for real clusters rather than for a demo.

Added

  • Enforcement survives an agent restart. The agent closed every BPF handle on exit, so a DaemonSet rolling update detached every program and the replacement started with empty maps: the node was unprotected for the whole reload and the learned baseline was gone. Programs, maps and links are now pinned to bpffs and adopted on startup. Adoption is refused when it would be unsafe, including a digest over the compiled BPF objects so new userland never runs against old programs, and any pinning failure degrades to the previous behaviour rather than refusing to start. Where pinned state cannot be adopted, allow-sets are rebuilt from ContainerProfiles, a path that is lossy by construction and documented as such.
  • Keyless cosign signatures on the released image, signed by digest rather than by tag, so moving a tag cannot leave an old signature still looking valid. Each release also carries an SPDX SBOM attested to the image, build provenance for the image and for install.yaml, and a SHA256SUMS signed with cosign sign-blob. make verify-release checks all of it and docs/packages.md documents the certificate identity to verify against. Signing is gated to tag pushes, so main and latest stay unsigned by design, and tags published before this release fail verification rather than passing quietly.
  • A PodDisruptionBudget for the operator, deliberately maxUnavailable rather than minAvailable: at one replica a minAvailable: 1 budget permits zero disruptions and makes the node permanently undrainable. The agent gets none by design, because kubectl drain deletes DaemonSet pods rather than evicting them, so a budget there protects nothing and wedges the descheduler once the node is cordoned.

Fixed

  • /readyz reported ready as soon as the health port bound, so a rollout advanced before any program was attached. It now reports actual attachment. It deliberately ignores the best-effort LSM hooks, so a kernel without BPF LSM does not block a rollout forever.
  • The agent loads and verifies its BPF programs before the manager binds its health port, so the only probe was a liveness check whose three default failures killed the container at roughly 55 seconds, on exactly the nodes where loading is slowest. Both components now have startup probes, with explicit timeouts and failure thresholds on the others.
  • The Helm chart shipped one of the three CRDs, and that one came from an older controller-gen. Helm never templates crds/, so helm install created no ContainerProfile and no AttackSurface, and the controllers watching them never synced. All three are now byte-identical to config/crd.
  • Chart RBAC was missing four grants that deploy/base has, so a Helm install crash-looped where kubectl apply of the same release did not, and every restart wiped the learned baseline. The chart pod spec is brought to parity with the base, including GOMAXPROCS, the namespace variable the telemetry join needs, and the seccomp settings.
  • Agent and operator resource requests sat below the BPF map footprint, which kernel 5.11 and later charge to the creating container's memcg, so a pod could be scheduled onto a node that could not hold its maps and be OOM-killed at map creation. Requests and limits are now sized against the ceiling the on-kernel test already enforces, and the guard reads that ceiling from the test rather than repeating the number.

3.4.1

2026-09-21

Changed

  • Batched four of five open Dependabot bumps onto one PR rather than merging them one at a time: github.com/onsi/gomega 1.43.0 -> 1.43.1, golang.org/x/net 0.58.0 -> 0.59.0, github.com/onsi/ginkgo/v2 2.32.2 -> 2.33.0, and github.com/yuin/goldmark 1.7.17 -> 1.8.6, plus their transitive updates. The fifth, google.golang.org/grpc 1.83.2 -> 1.84.0, was left out: govulncheck flags 1.84.0 with GO-2026-6443, a server panic via missing authority or Host headers, reachable through pkg/grpcapi's Serve call. The only fix is a v1.85.0-dev pseudo-version with no stable tag yet, so grpc stays at 1.83.2, which is itself a fixed version for the 1.83.x branch.

Fixed

  • Tee.Enqueue, which fans one event out to every live consumer on the export pipeline (the gRPC stream among them), had no test at all despite the documented guarantee that a refusal from one sink must not stop the event from reaching the others. Added coverage for the empty tee, all-accept, partial-refusal and a nil sink in the list, plus a benchmark for the hot path.

3.4.0

2026-09-19

The console, and the first answer to rare behaviour. Learning is a wall-clock window, so something a workload does once a day was never in its baseline and was refused like an attack. This release lets an operator declare it, and lets Pahlevan find a CronJob's schedule and learn for long enough to see it.

Added

  • An interactive console. Bare pahlevan in a terminal opens it: overview, policies, profiles, workloads, events, attack surface and coverage, reading the cluster and the agent's event stream. It is read-only by construction and cannot change a policy or a mode. Piped or redirected it prints help instead, so scripts keep working.
  • Declared expected behaviour. learningConfig.expectedBehavior lists files, destinations, executables and capabilities a workload uses rarely. They are seeded into the allow-set, additively and never wider than declared, and reported on the profile as declared rather than learned.
  • Cycle-aware learning windows. Pahlevan walks a pod's owners to its CronJob, reads the schedule, and raises the learning window to one full interval. A declared window is never lowered; the discovered one is capped at seven days, past which the answer is to declare the rare operation.
  • The v1beta1 API, now the storage version. v1alpha1 is still served, deprecated, and converts field by field; six paths that cannot round-trip are listed and tested.
  • Names for external destinations. Private ranges, CGNAT, link-local, the cluster CIDR and cloud metadata endpoints are labelled. Naming never blocks the event path, and reverse DNS is off by default.
  • Tracing that shows something: reconcile, the learning window, profile generation, and eBPF load and attach per program. The per-event path is deliberately untraced, and a test enforces it.
  • An optional dashboard, off by default and absent from install.yaml. Reads are authorised per request with a SubjectAccessReview; no writes, no cluster-admin, TLS required, no external origins.
  • Documentation published on the site, generated from docs/, with release articles generated from this file.
  • A scheduled check that opens an issue when a merged release was never tagged.

Changed

  • CI runs its checks in parallel: about 12 minutes to about 8 for a code change, and about 30 seconds for a documentation-only one.
  • The image builds from a 10 MB context instead of 15 GB, and a fresh VM provisions in 44 seconds instead of 88.
  • The documentation was rewritten against the code. It had described LSM hooks, policy fields, metrics, subcommands and error messages that do not exist.
  • The README is shorter and the architecture diagrams show all eight programs, in larger type.

Fixed

  • Every IPv6 destination was exported as 0.0.0.0.
  • --kubeconfig was passed as a context name, so it could silently select the wrong cluster. --context is now honoured too.
  • A replay lost the rest of a capture on one malformed line.
  • pahlevan ui > /dev/null was treated as a terminal.

Removed

  • Three pieces of tracing that implied a capability they did not have: a second tracer nothing read, an always-empty Traces field, and a tracer provider with no span processors that reported tracing as enabled.

3.3.3

2026-09-14

Changed

  • Batched three dependency bumps onto one PR rather than merging them one at a time: golang.org/x/term 0.45.0 -> 0.46.0, github.com/onsi/ginkgo/v2 2.32.1 -> 2.32.2, and sigs.k8s.io/controller-runtime 0.25.0 -> 0.25.1. controller-runtime 0.25.1 still declares go 1.26.0 and k8s.io/api v0.37.0, matching what this module already used, so the bump carried no transitive Kubernetes skew.
  • Unified three byte-for-byte identical label-selector matchers, one in each of PahlevanPolicyReconciler, ContainerLearnerReconciler and AttackSurfaceAnalyzerReconciler, into a single matchesSelector function. Only the first had a full test table; the other two sat at 16.7% and 33.3% coverage and could silently drift from each other on the next edit. The shared implementation now has one comprehensive table-driven test and a benchmark; internal/controller coverage moved from 74.6% to 79.1%.

3.3.2

2026-09-07

3.3.1 was merged but never tagged, so it was never published and nothing could install it. Its changes are released here.

Changed

  • Batched five open Dependabot GitHub Actions bumps onto one PR (actions/cache to v6, actions/configure-pages to v6, actions/upload-pages-artifact to v5, actions/deploy-pages to v5, softprops/action-gh-release to v3) rather than merging one at a time.

Fixed

  • The released install.yaml pinned nothing. It is the file attached to every release, and docs/packages.md points at releases/download/<version>/install.yaml as the immutable tag recommended for production. It shipped image: ghcr.io/obsernetics/pahlevan:latest, so that path pulled a tag that moves on every merge to main. The generator now stamps the release version into the manifest; deploy/base keeps :latest on purpose, because it is a kustomize base and an overlay sets its own tag.
  • Nothing regenerated or verified install.yaml in CI, so it could drift from the deploy/base and config/crd sources it claims to be generated from. hack/install now fails if it has drifted, if it does not pin the version the Makefile declares, if the kustomize base grows a hardcoded version, or if the docs send readers to a different release than the manifest deploys.
  • A data race in Manager.SetAction. It read the four eBPF collection pointer fields directly to decide which hooks to skip, with no lock held, while Load only ever writes them under the manager's mutex. A policy transition landing while the agent is loading or reloading its eBPF objects raced on those reads; confirmed with -race against a concurrent reload and fixed by having SetAction go through the four per-hook setters, which already take the read lock, rather than reading the fields itself. A regression test reproduces the race on purpose rather than relying on a scheduler to find it again.

3.3.0

2026-09-07

Added

  • A generic kprobe an operator points at any kernel function by name, with no rebuild. One pre-compiled program serves every probe: each attachment carries a bpf_get_attach_cookie holding the probe id, so forty probes are forty links over one program rather than forty copies of it. Up to five arguments are captured and up to four selectors are ANDed against them, or against the calling uid, gid or pid. Actions are report, audit, kill and signal - a kprobe fires alongside the function rather than in place of it, so it cannot refuse the call, and a policy asking it to is refused at load rather than quietly downgraded.
  • The kernel tests run in CI. They needed a VM with lsm=bpf and were run by hand, which meant a verifier rejection reached a human only if somebody remembered. Three did, in one week, and every one of them compiles, passes vet and passes the unit suite. A pull request touching bpf/, pkg/ebpf/ or the VM harness now boots a guest under KVM and loads all seven programs through a real verifier, and a nightly run catches a break that arrives from outside those paths. A test asserts the workflow's path filters cover every file a kernel decides about, so the job cannot go on reporting success by never running.

Changed

  • The event decode path is 1.67x faster and allocates 62% less: 3772ns and 68 allocations to 2255ns and 26 across the eleven decoders, measured by the benchmarks in pkg/ebpf. Fixed-width kernel fields were being turned into fresh strings per event, so a workload doing nothing unusual produced garbage proportional to its syscall rate. Process names and paths are now interned through bounded tables read without copying the bytes, and the per-event container id is derived rather than formatted. Argv is deliberately not interned: it is attacker-influenced, and an unbounded table keyed on it is a memory-growth primitive.
  • hack/vm/env.sh takes PAHLEVAN_VM_DISK, and up.sh checks /dev/kvm is usable before booting rather than failing inside qemu with an accelerator error buried in a serial log. The vm-test target honours PAHLEVAN_VM_CACHE instead of hardcoding .vmcache, which worked only for as long as nobody set the variable.
  • CI lints the workflows. A workflow with a malformed ${{ }} expression does not fail loudly: GitHub reports the run under the file's path instead of its name, ends it in zero seconds, and every job it was supposed to gate simply never appears - which looks exactly like a workflow that passed.
  • Three ATT&CK techniques the coverage table was missing, each evidenced by the entry's own description: T1552.001 (Credentials In Files) on lsm/file_open, the on-point technique for a per-path open monitor in a cluster where a service-account token is a file; T1548.001 (Setuid and Setgid) on kprobe/commit_creds, whose entry already explained that it separates a setuid binary's credential change from an unexplained one; and T1070.003 (Clear Command History) on uretprobe/readline, whose entry already named history -c as something it captures.

Fixed

  • A failing ring-buffer reader no longer spins. All eight readers shared a loop that, on any error other than a closed reader, continued immediately - so a persistently failing reader burned a core and filled the log at the rate it could produce errors. The eight copies are one function; a read error is counted on pahlevan_ebpf_read_errors_total, logged once, and backed off 100ms. The loop still exits promptly on stop.
  • The demo GIF claimed seven ATT&CK techniques Pahlevan does not list, and attributed two more to the wrong hook. The coverage table in the recording is typed by hand, because it is a recording, and nothing connected it to pkg/coverage. A technique shown in the GIF is a claim about what Pahlevan's data is evidence for, and the GIF is the first thing anyone sees. The recording now prints what pahlevan coverage prints, row for row, and a test fails if it ever shows a technique the table does not have, attributes one to the wrong hook, or drops a detector.
  • The CI formatting check split its file list on whitespace, so a path containing a space would have been checked as two files that do not exist. Found by the workflow linter this release adds.
  • A data race in the container tracker. Start runs discovery from two goroutines - the pod watch and the periodic refresh - and both reach updateContainer, which takes the write lock, so the map itself was safe. The container count logged at the end of discovery was read under no lock at all, racing with those writes. It surfaced on a CI runner under -race and would not reproduce in 150 runs on a workstation, so the regression test calls discovery concurrently on purpose rather than relying on a scheduler to find it again.

3.2.0

2026-09-06

Added

  • Formatted delivery of findings, closing the item that had been first on the Near term list since 3.0.0. A file, an HTTP webhook, OTLP and the gRPC stream all carry the event envelope, which is the right shape for a collector and the wrong shape for a person. Three sinks now carry formatted messages instead, all ordinary Exporters so they ride the existing bounded queue and drop counting:

    • Slack (--notify-slack-webhook): one Block Kit message per batch, with the fallback text set, because without it a push notification says "attachment".
    • PagerDuty (--notify-pagerduty-key, --notify-pagerduty-severity): one incident per distinct finding rather than per batch, keyed so the same problem re-triggers the same incident instead of opening another. The incident's source is the workload, not the node.
    • A Go template (--notify-template-url, --notify-template) for everything else. Parsed at startup, so a syntax error names the mistake rather than being logged once per batch while nothing arrives.

    All three send denials only unless --notify-all-events is set, and deduplicate per finding over --notify-dedupe-window, keyed on the owning workload rather than the pod so a rollout does not send one message per replica for a single misconfiguration.

Changed

  • Every direct dependency is current. The OpenTelemetry family to 1.46.0 (log and its exporters to 0.22.0), grpc to 1.83.2, testify to 1.12.1, prometheus/client_model to 0.6.3, ginkgo and gomega, and the Kubernetes stack to 0.37.0 with controller-runtime 0.25.0. Both k8s 0.37 and controller-runtime 0.25 declare the Go version the module already requires, so unlike the 0.36 move this needed no toolchain change and broke no API.

Fixed

  • The v3.1.0 release was merged but never tagged. The tag push failed and the release, its image tags and the chart never published. Tagged and verified.

3.1.0

2026-09-01

Added

  • pahlevan coverage, and the pkg/coverage package behind it: Pahlevan's own mapping of its seven eBPF programs to the MITRE ATT&CK techniques their observations can help an analyst confirm or rule out. Closes the "detection coverage vocabulary" gap ROADMAP.md has carried since 3.0.0 - previously the only way to answer "what does this cover" was reading bpf/*.c one file at a time.

Changed

  • pkg/discovery's ContainerTracker (the Kubernetes-facing container discovery used by the integration test harness) now has real unit test coverage: 30.4% to 94.9%, with a removed dead-code scaffold (ContainerScanner, PodInfo, NodeInfo, EventWatcher and friends in pkg/discovery/types.go) that was never wired into anything and only tested itself.
  • cmd/pahlevan-agent's event observer (agentObserver) and policy resolver's Refresh/PodMeta now have unit test coverage (26.8% to 54.7%); the remaining gap is main() itself, which loads real eBPF programs and cannot be unit-tested on the host.
  • cmd/pahlevan-operator's two admission runnables were extracted into testable functions (ensureAdmissionOnce, reconcileDerivedAdmissionOnce) with unit tests (0% to 15.3%); behavior is unchanged. The remaining gap is again main(), which requires a live Kubernetes API server.
  • docs/architecture.md's eBPF programs table was stale since the credential and shell monitors shipped: it listed five programs, including a lsm_monitor.c that does not exist, and was missing capability_monitor.c, cred_monitor.c and shell_monitor.c entirely. Now lists all seven with their real hooks.
  • docs/quick-start.md described a CRD shape and CLI that stopped existing release ago: learning/enforcement spec fields (real names are learningConfig/enforcementConfig), monitor/enforce mode strings (real values are Monitoring/Blocking), a pahlevan debug system-capabilities subcommand that was never implemented, a pahlevan_blocked_events metric that does not exist, and enforcement logs attributed to the operator when they are actually written by the agent DaemonSet. Corrected throughout.

Removed

  • docs/USAGE.md: unlinked from every other doc, and almost entirely fabricated (invented CLI subcommands like pahlevan debug system-info, pahlevan policy simulate, pahlevan policy apply -f policies/; a pahlevan-config ConfigMap tuning surface that was never built). Its accurate content already lives in docs/quick-start.md, docs/policy-reference.md and docs/deployment.md.
  • pkg/seccomp.Generate, a one-line wrapper around GenerateWithOverrides with no callers anywhere in the tree. GenerateWithOverrides(allowed, nil, nil) is the direct replacement for anyone calling it as a library.
  • internal/controller AttackSurfaceAnalyzerReconciler.syscallNumberToName, an unreachable duplicate of the syscall table pkg/ebpf/pkg/seccomp already maintain.

3.0.0

2026-08-31

Two hundred and twenty-five commits since 2.0.0. The theme, unintentionally, is honesty: a large share of this release is the discovery and removal of things that looked like they worked and did not. Where a feature could be implemented it was; where it could not, it now says so.

This is a major release for two reasons, both listed under Breaking changes: the gRPC event stream will not start unauthenticated any more, and the module now requires Go 1.26.

Breaking changes

  • The gRPC event stream refuses to start plaintext and unauthenticated.

    If you run the agent with --grpc-bind-address and no transport security, it now exits at startup instead of serving. The stream carries every denial on the node - which pods exist, which paths they read, which destinations they dial, the full command line of every exec - and that is a reconnaissance report.

    Previously it started anyway and logged a warning. The person who forgets the certificate and the person who reads the startup log are rarely the same person, so the warning was not a control.

    To upgrade, pick one:

    # mTLS, the right answer in a cluster. cert-manager can issue the pair.
    - --grpc-tls-cert=/etc/pahlevan/tls/tls.crt
    - --grpc-tls-key=/etc/pahlevan/tls/tls.key
    - --grpc-client-ca=/etc/pahlevan/tls/ca.crt
    
    # TLS plus a bearer token, for a collector that cannot present a certificate.
    - --grpc-tls-cert=/etc/pahlevan/tls/tls.crt
    - --grpc-tls-key=/etc/pahlevan/tls/tls.key
    - --grpc-token=$(PAHLEVAN_GRPC_TOKEN)
    
    # Or say explicitly that this listener is unreachable and you accept it.
    - --grpc-insecure
    

    A bearer token without TLS is still refused: in cleartext it is a token you have published. If you do not set --grpc-bind-address at all, nothing changes for you - the listener is off by default and always was.

  • The module requires Go 1.26.

    go.mod declares go 1.26.0, so anything importing github.com/obsernetics/pahlevan needs Go 1.26 or newer to build. This came from k8s.io/cli-runtime 0.36.3, which requires it, and which in turn forced controller-runtime to 0.24.1.

    Running the published container image is unaffected - it ships compiled binaries and needs no toolchain.

Added

  • Kernel-enforced processFilter. Constrains the parent process comm, the effective uid and the effective gid at bprm_check_security. The learned allow-set asks whether a container has ever run a binary; this asks whether the process running it is allowed to. That distinction is what covers the interpreter already in the image, which the allow-set cannot.
  • Container-breakout detection. An exec whose working directory belongs to a different mount namespace from the process is the invariant the runC breakout class violates (CVE-2024-21626 and its successors). Detected during learning as well as under enforcement, never added to the allow-set, and reported with its own flag, counter, alert and OTLP severity.
  • Destination naming. A denial reads prod/postgres:5432, resolved from Services, pods and nodes the agent already caches. An address the cluster does not know is tagged external, which separates a misconfiguration from exfiltration. No DNS query is made.
  • OTLP export for security events, using OpenTelemetry semantic-convention attribute names, plus one shared resource across metrics, traces and events so Grafana can join them. examples/observability/lgtm-stack.yaml deploys the collector and datasources.
  • pahlevan policy explain -f, which translates a policy offline and names every part the data plane will not enforce. --strict fails a CI gate.
  • gRPC TLS, mTLS and bearer-token authentication for the event stream, with the security posture printed at startup. The default is still plaintext.
  • allowDNS and allowLoopback are enforced, as a per-cgroup flag checked ahead of the allow-set.
  • Syscall arguments on every syscall event. The monitor moved from raw_tracepoint/sys_enter to tracepoint/raw_syscalls/sys_enter, whose format already carries the six arguments extracted, identically on both architectures. A watch set of escalation primitives - ptrace, unshare, setns, bpf, mount, io_uring_setup and the rest - bypasses the in-kernel deduplication, so they report every occurrence rather than only the first. The first ptrace a process makes is usually a debugger attaching at startup; the interesting one is the fourteenth.
  • Privilege-escalation detection at kprobe/commit_creds. Every other monitor watches a request; this one watches the result. commit_creds is the single function through which any task's credentials change, so a local-root exploit that overwrites a cred struct and calls it directly lands here having made no syscall at all. The discriminator is task->in_execve: a setuid binary gains privilege inside execve and sudo does it all day, while privilege gained with none underway has no other explanation. Needs no BPF LSM.
  • Interactive shell capture at uretprobe/readline. cd, export and history -c are shell builtins: they produce no exec, no open and no connect, so an exec-based monitor watches somebody work through them and reports nothing but the shell's own process. Attached per container through /proc/<pid>/root when a shell execs, and off by default behind --trace-shell-commands, because recording what a person types is a decision an operator makes deliberately.
  • Five enforcement actions instead of two. Learn, Deny, Kill, Audit and Signal, with a configurable errno and signal number, packed into one __u32 per cgroup so the hot path still costs one map lookup. Audit reports what would have been refused and refuses nothing, and deliberately does not learn - an audit pass that quietly added everything it reported would report each violation once and never again. Signal exists for SIGSTOP: freezing a process leaves its memory for an incident responder, where SIGKILL destroys exactly that.

Fixed

  • ContainerProfile status was silently discarded on every write. The resource has a status subresource, so a server-side apply to the main resource does not write status - a real API server drops it. Profiles were reaching the cluster with no counts, no learned syscalls, files, destinations or capabilities, no phase and no rollback history. The object existed and looked healthy. Found only when controller-runtime 0.24 started modelling the real server in its fake client.
  • Export formats returned canned strings. exportToMermaid returned the literal "graph TD" for every cluster; GraphQL and Cytoscape returned {}.
  • The metrics provider was built with no readers, so every recorded metric was silently discarded while --observability-exports reported the exporter as configured.
  • Nine metric recorders discarded their labels, so per-policy questions had no answer and the gauges were last-writer-wins across containers.
  • Event-handler errors were discarded, making a broken export pipeline indistinguishable from a quiet cluster.
  • An unbounded no-op event handler was registered on the ring-buffer hot path on every reconcile, racing on a stale pointer.
  • Observability shutdown was unbounded, hanging the agent on every rollout while a collector was down.
  • Four GaugeVecs were named _total, so rate() over them was nonsense.
  • mode: Off did the opposite of what it says - unquoted Off is a YAML 1.1 boolean, and there was no enum, so it silently became Monitoring. The field is now validated at admission.
  • pahlevan version required a Kubernetes cluster, and the Dockerfile injected build metadata into variables that do not exist, so every shipped binary reported an unknown commit and date.
  • Seccomp profiles were written world-readable into a world-readable directory.
  • Two goroutines ticking hourly to call empty functions; three unreachable vendor exporters guarded on fields nothing assigned; a test double compiled into every binary; and the two CLI columns that printed N/A over data one dereference away.

Changed

  • Dependencies: cilium/ebpf to 0.22.0, prometheus/client_golang to 1.24.1, go-logr to 1.4.4, the OpenTelemetry family to 1.45.0 (log and its exporters to 0.21.0), k8s.io/cli-runtime to 0.36.3 and controller-runtime to 0.24.1. otel/log 0.21 removed its own Value and KeyValue types in favour of the ones in go.opentelemetry.io/otel/attribute, so a log record's attributes are now the same type as a span's and a metric's.
  • The README's diagrams are rendered images, authored as HTML under docs/assets/diagrams/ and regenerated by hack/render-diagrams.sh, rather than ASCII art whose box borders did not line up.
  • Commit messages are checked by a hook. .githooks/commit-msg rejects assistant attribution trailers; scripts/setup-hooks.sh enables it.
  • Every policy example was invalid against the CRD and has been rewritten. About twenty-five invented keys were being silently pruned by the API server, so the examples applied cleanly and did a fraction of what they said. A strict-decode test now fails the build.
  • docs/api-reference.md is generated from the Go types. The hand-written version described an API that had never existed.
  • install.yaml, the Pages site and the demo GIF are all kept in step automatically - each had drifted, and each now has a check that fails when it does.
  • Seven CRD fields are documented as inert rather than left to look functional.
  • make vm-test runs every VM test. It filtered on TestVMLoad, so nineteen tests written for the VM had never run there.

2.0.0

2026-08-14

Pahlevan 2.0.0 is a redesign. The single all-in-one operator is replaced by a privileged per-node agent plus an unprivileged, leader-elected control plane, and the eBPF data plane moved from placeholder code to real CO-RE programs that observe and deny in the kernel.

Added

  • Two-workload architecture: pahlevan-agent (privileged DaemonSet, owns the eBPF data plane) and pahlevan-operator (leader-elected Deployment, no host access, runs with hostUsers: false).
  • CO-RE syscall monitor: raw_tracepoint/sys_enter with a ring buffer and in-kernel deduplication per (cgroup, syscall).
  • CO-RE file monitor on lsm/file_open with path resolution via bpf_d_path() and graceful degradation when the BPF LSM is unavailable.
  • In-kernel file enforcement: unlearned opens are denied with EPERM.
  • In-kernel network egress enforcement via lsm/socket_connect, preceded by a CO-RE kprobe/tcp_connect network monitor for observation.
  • Process/exec monitoring and enforcement via lsm/bprm_check_security.
  • Adaptive learn to enforce loop: per-cgroup allow-sets are built during the learning window and the agent transitions to enforcement autonomously.
  • Seccomp profile generation from the learned syscall set.
  • CEL ValidatingAdmissionPolicy hardening for PahlevanPolicy resources.
  • ContainerProfile CRD with profile persistence, a metrics endpoint, and an enforcement counter.
  • AttackSurface CRD for cluster-wide posture aggregation.
  • Container attribution: cgroup id to Kubernetes pod/container resolver.
  • hack/vm/: reproducible QEMU/KVM harness that provisions a kernel with the BPF LSM enabled for eBPF load, attach, observe, and enforce tests.
  • Benchmark harness measuring detection, prevention and overhead, with a no-agent control pass (test/benchmark/, docs/benchmarks/).
  • GitHub Pages site (pages/) with a deploy workflow, published Helm chart, install and benchmark sections.
  • Committed bpf2go bindings and eBPF objects so the tree builds from a clean checkout without clang on the build host.
  • Documentation and demo: README with badges, an animated demo GIF, and an architecture diagram.

Changed

  • Helm chart and install.yaml rewritten for the agent plus operator split.
  • Build produces three binaries: pahlevan-agent, pahlevan-operator, and the pahlevan CLI; the container image is a minimal distroless runtime.
  • Go toolchain moved to 1.25 across go.mod, the Makefile, and the Dockerfile.
  • Dependency bumps, including Kubernetes libraries to 0.35.0, sigs.k8s.io/controller-runtime to 0.22.4, github.com/cilium/ebpf to 0.20.0, github.com/spf13/cobra to 1.10.1, Ginkgo/Gomega, OpenTelemetry, go-logr, and prometheus/client_golang.
  • Substantially expanded test coverage across the controller, learner, policies, CLI, APIs, observability, metrics, attribution, and visualization packages, plus decode round-trip tests and benchmarks for the eBPF event parsers.
  • CI restructured into a build, push, then test flow with dependency caching; security workflows consolidated.
  • Documentation cleaned of em dashes project-wide and the architecture SVG redrawn with clean connectors.

Fixed

  • Agent crash loop caused by a double eBPF attach and a duplicate controller name.
  • Syscall monitor silently disabled by default because of ARRAY map zero initialization.
  • Panic in the metrics path.
  • go.sum inconsistency from an unused golang.org/x/exp dependency.
  • Site layout overflow and horizontal scrolling in long code blocks.
  • CodeQL and vulnerability workflows failing because eBPF bindings were not generated before the build.
  • Reachable CVEs patched; govulncheck clean.

Removed

  • Admission webhook package and its integration wiring, replaced by the CEL ValidatingAdmissionPolicy.
  • Dead code: the pkg/events package, an orphaned eBPF mock.go, and an unused eBPF manager.
  • Placeholder and stubbed subsystems, replaced by working implementations.

1.0.0

2025-09-25

Added

  • First public release of Pahlevan: an eBPF-based Kubernetes security operator with the PahlevanPolicy CRD, a learning phase, enforcement modes, self-healing, observability, and Helm plus manifest based installation.

Ready to install?

Apply the manifest, label a workload, and let the learning window do the rest.