外部探针、追踪查询、告警响应
ARCHITECTURE DOSSIER · 07
日志、可观测与恢复
Nginx、systemd、Docker、Loki、Grafana、备份与故障演练
OPSCODE + RUNTIME + QArequest_id、route、status、latency
structured event、domain ids 与 errors
Docker/journal/Nginx 采集、14d 索引、Explore 与 4 个 dashboard
最小 Docker discovery API;stream coverage 仍需独立证明
组件状态、影响、时间线、决策与跨域证据
当前 dump/tar/local+S3;目标 manifest、隔离恢复与 attestation
业务/日志/本地副本共根卷;MariaDB/Gitea archive 有 off-host copy
CRITICAL SEQUENCE
真实主链与事务边界
每一步都必须能回到调用方、契约、状态 owner 与失败证据;UI 只显示服务端已经确认的真值。
- 01
外部探针、Nginx、systemd、Docker readiness 分别产生带 as_of 的 HealthSignal
→ - 02
服务写入安全结构日志并绑定 request/run/entity/release 字段;高基数值不做 label
→ - 03
Alloy 当前采集 Docker、Nginx JSON access 与持久/运行 journal,写入单节点 Loki
→ - 04
Log Explorer 以时间/环境/service/source 为预算边界查询,并在 detail 关联 request/run/entity
→ - 05
异常只有在 impact evidence 形成后进入 Incident;时间线区分信号/事实/假设/决定/动作
→ - 06
审计 projection 关联 actor/session/resource/policy/domain guard/receipt,但原始事件仍归领域 owner
→ - 07
每个数据 owner 生成 BackupManifest 与 off-host copy,在隔离环境恢复并执行真实服务旅程
→ - 08
记录 actual RPO/RTO、失败检查、双人签署与 cleanup receipt 后才算 recovery pass
INTERFACE REGISTER
跨项目接口与安全条件
producer、consumer、契约和 guard 同时登记;其中任何一列缺失都不能进入多团队并行开发。
HealthSignal {component,source,state,as_of,detail}Status / Incident targetsignal != product availability; signed snapshot
AccessEvent {request_id,host,path,status,duration}Alloy/Lokiquery/token/PII redaction before ingest
StructuredLogEvent {level,event,request,run,entity,release}Alloy/Lokino secrets; bounded labels; tenant-aware query
AuditEvent {actor,resource,action,decision,receipt,hash}Federated audit detailappend-only owner source + safe diff
IncidentTimelineEvent {kind,fact,evidence,owner,at}Status / Postmortemverified updates + immutable history
BackupManifest {domains,as_of,artifacts,hash,target,tool_version}Restore drilloff-host copy + isolation + two-person approval
STATE OWNERSHIP
数据真值与恢复
STOREjournald + Loki filesystem on root EBS
RETENTIONLoki 14d;journal 3.2 GiB observed
RECOVERY诊断数据,不是业务恢复源;当前无 off-host log copy
STOREappend-only event/projection TARGET
RETENTION按事件/审计类别
RECOVERY从 owner events 重建并验证 hash
STORElocal + S3 current
RETENTIONlocal count policy; S3 policy not audited
RECOVERYgzip verified only;隔离 restore 未证明
STOREhost/Docker volumes current
RETENTION不统一
RECOVERY当前主备份 manifest 不完整
FAILURE MODEL
故障必须怎样收敛
SYMPTOM业务、日志与本地副本同时丢失
CONTAINMariaDB/Gitea archive 有 S3 copy;其余域与诊断证据仍不完整
domain inventory + P0 recovery gapSYMPTOM部分容器日志可能缺失但 Loki/Alloy 表面仍运行
CONTAIN比较期望容器与实际 stream;proxy healthy 也不能跳过 coverage canary
proxy health + expected/observed stream coverageSYMPTOMLoki 写入/查询退化
CONTAIN高基数字段放 body,不做 label
ingestion metricsSYMPTOM凭证或 PII 泄漏并被 14d 保留
CONTAINingest 前 redaction + deny tests + safe export
secret/PII scanSYMPTOM灾难时无法启动服务
CONTAIN以 restore drill 作为唯一通过标准
drill report + RTOQA RELEASE GATES
不是画完架构图就算完成
这些门禁连接 UI 状态、接口负向路径、运行证据与可恢复性;发布证据必须绑定 exact commit。
request/run/entity/release 跨 edge/service 可关联
synthetic trace + field coverage
日志/trace/export 不含 token、signed URL、DSN 与 PII
ingest redaction + negative scan
期望服务/容器与日志 stream 覆盖一致
discovery coverage gate
HealthSignal、产品影响与状态更新不混为一谈
status/incident contract test
Incident timeline 事实、决定、动作均带 owner/evidence
exercise + postmortem gate
Audit actor/resource/decision/safe diff/hash 完整
cross-domain audit fixture
BackupManifest 覆盖全部责任域与 off-host hash
P0 inventory gate
隔离环境恢复服务旅程并记录 actual RPO/RTO
季度 two-person restore drill