feat: add deployment foundation and cross-device handoff
This commit is contained in:
@@ -0,0 +1,79 @@
|
||||
# Operation metadata store
|
||||
|
||||
Internal execution foundation; not a CLI write endpoint, application executor,
|
||||
backup engine or systemd supervisor. The existing CLI remains read-only.
|
||||
|
||||
## Contract
|
||||
|
||||
Bootstrap must supply one fixed, existing, administrator-owned local directory
|
||||
per host. Use the same canonical directory for every runner on that host.
|
||||
Untrusted users/applications must not be able to replace the directory, its
|
||||
ancestors or files. Linux deployment permissions should be 0700 for the directory
|
||||
and 0600 for metadata. Windows ACL configuration belongs to the installer.
|
||||
NFS/SMB/distributed locking is not supported.
|
||||
|
||||
`Acquire(directory, hostID)` opens an os.Root and takes a nonblocking exclusive
|
||||
OS lock on `host.lock`. Different paths/roots are not a distributed host registry.
|
||||
Hold the returned Session throughout any future application write operation.
|
||||
Never remove/replace `host.lock`, including during cleanup: its inode is the lock
|
||||
identity. Close releases it; abrupt process termination releases it at OS level.
|
||||
|
||||
- `Begin(id, planHash)` registers queued work and returns `(operation, created)`.
|
||||
Reusing ID/hash returns the original record with created=false. Another hash
|
||||
conflicts. Different IDs are blocked while an unresolved record exists.
|
||||
- `Advance(id, revision, next)` compares revision and validates the transition.
|
||||
A queued-to-running transition is the claim; a second claim fails.
|
||||
- `Get(id)` returns a value copy. Closing or poisoning a session forbids its use.
|
||||
|
||||
Allowed transitions:
|
||||
|
||||
```text
|
||||
queued → running | cancelled
|
||||
running → succeeded | failed_recovered | needs_attention | unknown
|
||||
needs_attention / unknown → succeeded | failed_recovered
|
||||
terminal states → no transitions
|
||||
```
|
||||
|
||||
Success/recovery labels are assertions by the caller, not proof. The future
|
||||
executor must verify actual effects before recording them. Unknown/running work
|
||||
survives restart without replay. Reconciliation requires inspecting actual state;
|
||||
no automatic retry, forced reset, lease expiry or stale-lock deletion is provided.
|
||||
|
||||
## Persistence
|
||||
|
||||
state.json is a versioned, canonical JSON snapshot, maximum 16 MiB, with host
|
||||
identity and SHA-256 integrity checksum. Unknown fields, duplicates, truncation,
|
||||
changed data and host mismatch are rejected. The checksum is not authentication.
|
||||
It is not an append-only audit log; event logging is a separate future component.
|
||||
|
||||
Mutation writes a unique private temporary file, syncs it, closes it, renames over
|
||||
the snapshot and (Linux) syncs the containing directory. A save error poisons the
|
||||
session because the rename may already have happened; close, reacquire and query
|
||||
before deciding anything. Temporary remnants from a killed process are ignored,
|
||||
never interpreted as successful work; automatic cleanup is not implemented.
|
||||
|
||||
Current snapshots retain all operation IDs. No pruning is provided, because
|
||||
forgetting completed IDs can re-enable an old request. At the size limit, writes
|
||||
fail closed. Admission reserves 64 bytes for the active record's later status
|
||||
and revision growth; capacity rejection does not poison read access. A retention/tombstone design is required before bounded production
|
||||
history cleanup is introduced.
|
||||
|
||||
host.lock also contains a synced initialization marker. Once work is persisted,
|
||||
a missing snapshot is rejected rather than treated as a fresh store. A crash
|
||||
between the first snapshot and marker can be repaired only from a valid snapshot.
|
||||
Protect both files. Restoring an older valid snapshot still rolls back idempotency
|
||||
history; do not restart execution without independent reconciliation. This package
|
||||
cannot detect malicious administrator edits or rollback/deletion of the entire
|
||||
state directory.
|
||||
|
||||
## Platform boundary
|
||||
|
||||
Linux uses flock and file/directory fsync. Windows uses LockFileEx for development
|
||||
tests, file sync and rename; Windows power-loss durability is not promised.
|
||||
Other OSes refuse acquisition. The local macOS panel will communicate with the
|
||||
Linux runner, not use this package as its local SQLite replacement.
|
||||
|
||||
Kernel locks are advisory on Linux; all participating writers must obey them.
|
||||
External Docker/Portainer operations are outside this lock and require drift
|
||||
checks. Multi-process tests prove process-crash behavior, not sudden power failure
|
||||
or storage-hardware reliability.
|
||||
Reference in New Issue
Block a user