Nikhil Sood / Case studies

Runtime engineering case study / 7 min read / Updated

Building Nikxius Runtime: Execution Control and Recovery

I built Nikxius Runtime to carry one supported operation from authorization through execution and recovery inside a customer-controlled environment. The work now includes exact action contracts, a durable PostgreSQL lifecycle, Kubernetes rollback, an operator console, portable evidence, and a design-partner onboarding path.

  • 6 implemented operation profiles
  • 232 Runtime tests in the 12 September record
  • 27 native Kubernetes cases included in that record
  • 10 rollback acceptance assertions passed on 15 September

From payment evidence to execution control

Nikxius started with a payment failure window: a provider could accept an action before the local application durably recorded it. A timeout could then tempt the application to submit a second payment. The original research examined intent-before-effect persistence, human approval bound to exact bytes, and evidence that could be verified outside the issuing service.

The current Runtime applies that engineering focus to supported production automation. The product question is concrete: which exact operation may this agent request, what must still be true when it executes, and how will an operator determine what happened if the reply disappears? Payment controls remain a reference implementation; execution control and recovery are now the main product.

A customer-controlled Runtime

The Runtime accepts typed proposals from an agent or conventional automation. A bounded task grant names permitted effects, targets, and expiry. Identity alone is insufficient. Policy, native preconditions, and any required human review must all permit the exact command before dispatch.

The customer-hosted Node holds the constrained native resource credential. It reads native state independently, freezes the command and its bindings, and persists dispatch identity in PostgreSQL before making the native API call. The protected agent must lack an alternate write path. Existing identity, Kubernetes RBAC, admission policy, and monitoring remain part of the boundary.

The operation record separates native commitment from observed convergence. A resource change can be committed while the application is degraded or unobserved. The operator console exposes operations, approvals, uncertainty, reconciliation, and evidence without turning a resolved record into an automatic claim of success.

Six explicit operation profiles

I implemented a small set of typed contracts with operation-specific native checks and recovery rules. Support is limited to enrolled targets and declared effects. These profiles do not grant arbitrary shell access or control every tool an agent can call.

  • Kubernetes Deployment image change and explicitly linked image rollback.
  • Bounded scaling for stateless Deployments.
  • Selected non-secret ConfigMap key updates.
  • CronJob suspension.
  • Confirmation of a saved HCP Terraform plan. HCP provider behavior has mocked test coverage; real-account conformance remains unvalidated.

UNKNOWN is an outcome to preserve

A lost reply does not prove that a write failed. After dispatch has been recorded, an ambiguous result stays UNKNOWN. The Runtime retains the original operation, attempt, and reservation through restart, then reconciles available native evidence without blindly submitting another mutation.

A matching current image is insufficient evidence on its own: another writer could have produced that state. Recovery needs observations that can be tied to the original operation and its native resource history. If that history is missing, UNKNOWN can remain unresolved. The system makes that limitation visible instead of manufacturing certainty.

I kept commitment, observation, and integrity separate because recovery decisions depend on them. A committed change is not proof of application health, and a detected mismatch is not proof that nothing happened. Any compensating action needs its own authorization and outcome.

An exact Kubernetes rollback

The first design-partner workflow is kubernetes.deployment.rollback_image.v1: restore one explicitly permitted prior immutable image for one enrolled Deployment container. It must refer to an earlier committed Runtime image operation in the same tenant, on the same native target and UID, with the required reversed before-and-after image relationship.

Rollback is a new operation with its own identity, grant, review when required, attempt, and result. It is not a generic undo of any historical Deployment version, and its own response can be lost. Restoring an image does not restore database schemas, dependencies, data, or application health.

The public walkthrough remains a recorded kubernetes.deployment.set_image.v1 image-change test. The rollback contract has separate acceptance evidence. Keeping those demonstrations distinct prevents a polished replay from claiming a deployment it never performed.

Verification that exercises failure

The Runtime verification compiled on 12 September 2026 records 232 distinct passing Runtime tests, including 27 native Kubernetes cases. The native environment was disposable Kubernetes 1.35.0 with synthetic identity and MFA fixtures. The real PostgreSQL restart test and mocked HCP behavior have separate boundaries; these are not customer production results.

The 15 September rollback acceptance run passed ten assertions, exercising an actual rollback, response loss, Runtime process termination with SIGKILL, recovery, and retained uncertainty. A separate PostgreSQL restart test used a typed native-provider double. The two tests must not be described as a simultaneous Kubernetes-and-database failure experiment.

I also exercised native Kubernetes controls as a comparison: RBAC, selected admission rules, conditional patches, and retained observations. Those component checks establish overlap with the native stack. They do not establish lower customer cost or superior reliability for Nikxius.

The onboarding is part of the product

An implementation is difficult to evaluate if another engineer cannot install it, inspect its authority boundary, and operate it. I built prerequisite and readiness checks, redacted enrollment inspection, rollback draft preparation, and fixed file-based Runtime API commands. The client checks local output conflicts before a mutation and does not automatically retry a missing mutation acknowledgement.

The evaluation path covers one named nonproduction environment: establish identities and the authoritative writer, enroll the exact target, review the permitted prior image, issue the grant, propose and approve the operation, inspect its outcome, reconcile uncertainty, and export signed evidence. Kickoff, security review, native comparison, acceptance, and removal have documented owners and checks.

The public guide can be read in the browser or downloaded as PDFs. Customer source delivery uses a reviewed selective export rather than the full working repository. Later internal rehearsals on 16 September exercised fresh setup, a manual controlled rollback, verified export, response loss, and restart. Independent customer completion and customer support effort remain unobserved.

Portable evidence remains useful

The original payment-control work produced the NXC1 certificate format, hash-chained ledger evidence, public verification specifications, and a zero-dependency offline verifier. The financial v3 reviewer remains available as a separate browser-local reference tool. Its sample evidence is distinct from Runtime operation records.

Verification checks supported cryptographic bindings relative to the supplied trust keys. It can detect altered signed material, but an admitted signer can still make a false assertion. It does not establish bank settlement, universal mediation, sound human judgment, regulatory compliance, or application health. The three original papers retain those trust limits alongside the newer execution-control research.

The next proof is a customer-operated workflow

The current evaluation is bounded to one supported Kubernetes rollback in one nonproduction environment. A real team still needs to establish identity integration, source access, native credential exclusivity, alternate-writer review, security and retention ownership, and its own acceptance criteria.

I have built and internally tested the product and onboarding path. Customer independence, production admission, repeatability across customer environments, and a measured advantage over the native stack remain separate questions. Those are the next things to learn through the design-partner evaluation.

Sources and further reading

  1. [1] Nikxius Runtime architecture. Customer-controlled execution and the native credential boundary.
  2. [2] Action contracts. Exact authority, command bindings, native state, and required review.
  3. [3] Recovery model. Commitment, observation, reconciliation, and retained UNKNOWN.
  4. [4] Kubernetes rollback workflow. The permitted prior-image contract and earlier-operation requirement.
  5. [5] Historical Runtime verification. The dated 12 September local record; not customer production certification.
  6. [6] Design-partner evaluation guide. Current public evaluation materials, internal validation boundaries, and customer prerequisites.
  7. [7] When AI Gets Write Access. The newer research on authority, native execution, and recovery.
  8. [8] Effect-Before-Persistence in Automated Payouts. The original persistence failure window and payment recovery model.
  9. [9] Cryptographic Evidence of Human Oversight. The original digest-bound maker-checker evidence model.
  10. [10] Portable Evidence and Offline Verification. Certificate integrity, offline verification, and residual trust assumptions.
  11. [11] Financial v3 evidence reviewer. The separate browser-local payment evidence reference tool; sample evidence is not a live payment.