# User and AI Roles and Responsibilities Policy

Status: accepted P1-0 field-level draft; Sibyla-scoped policy. Source draft:
`Specs/Roles_and_Responsabilities_Policy.txt` in `invoice-skill-build` detached at
`9359c67c4ef0101218d7e0ffff1986114ba5cc7a`. Adapted only for the deterministic .NET production
harness; no prototype data is included.

## Purpose and precedence

Sibyla is autonomous by default within approved Rules of Engagement. This policy sits above and
supersedes Sibyla procedures where they conflict. It does not supersede or amend the FDR
prototype's thirteen binding documents: FDR retains its existing AI-execution role until and unless
that system separately migrates. A change affecting this Sibyla policy requires the Policy Change
Validation pre-flight before implementation; no FDR-side pre-flight is required for Sibyla
adoption.

## Core principle

Users own business rules and governance. Claude/AI reads and transcribes evidence and proposes
classifications, matches, corrections, and reusable improvements. It never captures, decides,
approves, persists, changes state, issues a code, writes or moves a file, archives, or calls an
integration. Authenticated deterministic .NET code alone validates and executes every approved
capture, decision application, state change, code issuance, database write, archive action, and
external integration action.

Human review is required when evidence is insufficient, materially conflicting or ambiguous, when
an outcome may materially affect finance, tax, compliance, master data, or policy, or when an
engaged human requests review. Human engagement outranks an automatic ItemClass assignment.

## User responsibilities

- Define, approve, refine, and retire Rules of Engagement.
- Supply authorized business context and ground truth through PAYCTR/PAYDTL, RCVCTR/RCVDTL,
  review decisions, observations, and approved master-data tables.
- Resolve novel, material, or weak-evidence situations.
- Approve changes that materially alter financial outcomes, classifications, reconciliation
  logic, policies, or controls.
- Decide whether a proposed learning becomes a reusable rule, mapping, alias, or match type.

Users do not have to process every routine record. Manual actions remain authenticated, scoped,
reasoned, and audited.

Disposition authority is separated. A reviewer may discard a record that has no final business
disposition. Discarding a `Posted` or `ReferenceOnly` record requires a request by one authenticated
actor and a counter-signature by a different actor holding `DispositionCounterSign`; no actor may
approve their own request. Purge remains confined to the configured purge role and its stricter
historical eligibility rule. `RestoreForReview` is an audited reviewer command whose exact target
is disposition `NULL` plus `DocumentStatus.AwaitingReview`; it never silently reinstates a prior
`Posted` or `ReferenceOnly` result.

The legal entry states use the existing `DocumentStatus` literals, not new generic status phrases:
a fresh capture starts with disposition `NULL` plus `DocumentStatus.Registered`; a malformed
response remains disposition `NULL` plus `DocumentStatus.RetryScheduled` while retryable, or
`DocumentStatus.NeedsAttention` when retry is not scheduled or has been exhausted; and a
schema-valid `NOT_A_DOCUMENT` remains disposition `NULL` plus `DocumentStatus.AwaitingReview`.
These operational states do not alter the persistence distinction: malformed output has an
ExtractionAttempt and DOCFAI but no ExtractionRevision or extraction-derived DOCLOG, while a
schema-valid `NOT_A_DOCUMENT` follows the valid-response path through ExtractionRevision, DOCLOG
and DOCFAI. The existing authenticated, counter-signed disposition governance remains unchanged.

## Claude/AI proposal responsibilities

- Read staged evidence and transcribe only what that evidence supports.
- Propose mappings, classifications, matches, corrections, and reusable improvements under the
  approved Rules of Engagement.
- Use inference only where policy permits it and evidence is sufficient, and label uncertainty.
- Detect conflicts, weak evidence, control breaches, and genuinely decision-requiring exceptions.
- Propose reusable improvements and test generalization against comparable cases.

Claude/AI output is inert proposal data. It cannot authenticate an actor, approve its own proposal,
perform a capture, apply a decision, mutate a queue or business record, allocate a permanent code,
write an archive object, or invoke a bank, ERP, storage, messaging, or other integration.

## Authenticated deterministic-system responsibilities

- Resolve the actor and company from authenticated server-side context, then capture and register
  documents and official bank extracts under approved procedures.
- Validate proposals and apply approved mappings, gates, classifications, and reconciliation logic.
- Perform approved captures, decision application, state transitions, permanent-code issuance,
  persistence, archive/file operations, and external integration calls idempotently and audibly.
- Reassess relevant open and historical records when an approved rule changes without overwriting
  human decisions or retroactively blocking grandfathered history.
- Preserve source evidence, rule/version, confidence, risk, decision, actor, correlation identity,
  and applied result; fail closed on missing authority, evidence, or company scope.

## Autonomous operation and escalation

Deterministic code may apply an already-approved rule automatically when it has strong evidence,
an approved recurring pattern is conflict-free, or a low-risk inference is explicitly permitted.
An AI proposal never supplies execution authority. Deterministic code creates a Decision-class
review item only when a human answer could change the outcome: competing evidence, missing
essential information, an unsafe new pattern, material financial/tax/compliance effect, or
conflict with an approved rule or decision.

DOCRQE and RECREV are exception-and-learning surfaces, not mandatory work queues. In DOCRQE, only
Decision-class items may be Open or InReview; the generated annex governs the richer terminal
status vocabulary, so Status and Annotation items are not forced to `Recorded`. RECREV has no
pinned ItemClass or `Recorded` status and remains governed by its own generated status vocabulary;
its Open items are human-engaged reconciliation decisions. Only DOCRQE Decision items contribute
to the DOCRQE “percent worked” denominator. An authenticated human request always creates or keeps
the relevant Decision work visible regardless of an earlier automatic classification.

## Learning from approved input

For approved user input, Claude/AI may identify comparable records and propose a reusable artifact.
Authenticated deterministic code must validate the cited record and authority, apply the approved
input, check conflicts and evidence, create any approved controlled artifact, re-evaluate comparable
records only when broader use is approved, apply it prospectively, and retain the evidence and
reason. One decision is never generalized when it may be exceptional or weak. A rejected proposal
becomes a durable matcher constraint and is never proposed again.

## Data-entry boundaries

- BNKMOV originates exclusively from official bank extracts or an approved bank integration.
  AI and users cannot invent or alter movements to make reconciliation work.
- Manual financial entries are allowed only in FDCHDR/FDCDTL under the Financial Document Entry
  Policy, with actor, reason, evidence, and audit.
- Payment/receivable input belongs in PAYCTR/PAYDTL and RCVCTR/RCVDTL. It is ground truth for
  matching but not proof of bank settlement without an official BNKMOV link.
- A wrong-but-plausible value is worse than a blank. Missing identifiers, rates, or evidence fail
  closed to review and are never estimated.

## Controls, audit, and improvement

Every automated result, learned rule, and review decision is traceable to its source, evidence,
rule or heuristic, confidence/risk, user approval where required, target record, AppliedReference,
timestamp, process identity, and reviewer. Recorded decisions survive full reruns. New rules are
prospective. A superseding state change closes its predecessor and either identifies its successor
or records a controlled reason why none exists; it is never blank.

Any change to a confidence, consistency, risk, or autonomous-action threshold requires the Policy
Change Validation pre-flight. This includes BNKMAT `MaxDateWindowDays` and `ToleranceAmount`,
DOCEFL `RiskFactor` and `ReviewPriority`, FX tolerances, and exact-match bands. Duplicate-payment or
duplicate-settlement risk, balances that conflict with evidence, and any attempt to change an
approved closed result are named review triggers; duplicate settlement always creates the
Mandatory RECREV review and applicable Block Reconciliation finding before execution.

After every three rounds of material change, review queue composition, recurring corrections,
conflicts, metrics, and control outcomes. Improvements must increase accurate, autonomous,
auditable operation without weakening source integrity, duplicate prevention, or financial
controls. A metric improving because a check was removed is a regression.
