Daewon Desk / Implementation evidence

Prototype checks.

A record of what we tested in the local policy-review sandbox, including the limits of those checks.

Recorded 11 October 202617 automated tests passedTargeted browser checks completed

Rules and review state, tested locally.

The recorded automated run completed with 17 passing tests and no failures. The suite exercises the sandbox's policy rules, source construction, input validation, and review state using fictional policy data.

These are software checks for a deterministic prototype. The sandbox does not call Claude. Passing these tests does not measure language-model accuracy, general document retrieval, resistance to arbitrary instruction injection, or real customer outcomes.

The browser observations below cover selected interactions in the same prototype. They are not a claim of an exhaustive accessibility, security, or browser-compatibility audit.

What the 17 tests cover.

Grouped coverage of the existing test suite
AreaChecked behavior
Policy changesEdited dispatch ranges, return windows, conditions, and exclusions appear in the relevant templated draft and evidence.
EvidenceReady-result evidence matches the current generated source set. Korean and English drafts use the same policy values and source IDs.
Missing informationRequired shipping or return-policy omissions withhold a draft. The engraving example produces no invented evidence.
Conflicting valuesDifferent return windows in two sources block drafting, including in the ordinary return case. Reconciling the values restores drafting.
Input boundariesInvalid numbers, reversed ranges, unknown options, blank or overlong messages, and unknown case IDs are rejected. Freeform message text is not used as a policy source.
Fixed untrusted casesThe designated untrusted example and specific instruction patterns withhold a draft. This is a finite rule check, not a general injection defense.
Review and exportExport requires review. Input changes, draft edits, new case or language results, and reset revoke prior approval. Blocked, stale, empty, or overlong drafts cannot be approved or exported.

The full test file below contains the individual assertions and values. It tests local rules and state transitions; it does not evaluate the meaning of arbitrary customer questions.

Observed in the browser.

  1. Changing a policy changes the result. An edited shipping value appeared in both the draft and its displayed evidence after rebuilding.
  2. Missing information withholds the draft. The unsupported example showed a hold state instead of an invented answer.
  3. A conflict remains visible. Return windows of 7 and 14 days blocked drafting and exposed the conflicting policy values.
  4. Resolving the conflict updates the Korean reply. Changing both return windows to 14 days restored the draft, with 14 days reflected in Korean.
  5. Review controls the text export. A draft required review before its local TXT export. The downloaded result contained the reviewed draft and supporting evidence.
  6. An edit revokes approval. Editing a reviewed draft required a new review before export became available again.

These observations establish that selected controls affect the underlying draft and review state. They do not establish speed improvements, translation quality across unseen text, or readiness for customer records.

Run the same rule checks.

Download these three files into one folder. The published test copy uses relative imports; its test cases and assertions match the project suite. A Node.js runtime with the built-in test runner is required.

node --test prototype-tests.mjs

The original command from the project workspace:

node --test work/daewon-tests/engine.test.mjs

Re-running the suite checks the files you downloaded. The recorded result above is dated; future code changes should be followed by a new run. Browser interactions and download contents require separate browser checks.

Live-model evaluation is still ahead.

The next proposed implementation adds approved-document retrieval and server-side Claude drafting. That work needs a separate evaluation of factual claim support, policy constraints, citation support, correct escalation, bilingual quality, latency, and cost.

A consented merchant pilot is also planned, not completed. The product brief describes the evaluation approach and a target of three shops over two weeks. No model-performance or customer-outcome results are reported here.