Daewon Desk / Implementation evidence
Prototype checks.
A record of what we tested in the local policy-review sandbox, including the limits of those checks.
Rules and review state, tested locally.
The recorded automated run completed with 17 passing tests and no failures. The suite exercises the sandbox's policy rules, source construction, input validation, and review state using fictional policy data.
These are software checks for a deterministic prototype. The sandbox does not call Claude. Passing these tests does not measure language-model accuracy, general document retrieval, resistance to arbitrary instruction injection, or real customer outcomes.
The browser observations below cover selected interactions in the same prototype. They are not a claim of an exhaustive accessibility, security, or browser-compatibility audit.
What the 17 tests cover.
| Area | Checked behavior |
|---|---|
| Policy changes | Edited dispatch ranges, return windows, conditions, and exclusions appear in the relevant templated draft and evidence. |
| Evidence | Ready-result evidence matches the current generated source set. Korean and English drafts use the same policy values and source IDs. |
| Missing information | Required shipping or return-policy omissions withhold a draft. The engraving example produces no invented evidence. |
| Conflicting values | Different return windows in two sources block drafting, including in the ordinary return case. Reconciling the values restores drafting. |
| Input boundaries | Invalid numbers, reversed ranges, unknown options, blank or overlong messages, and unknown case IDs are rejected. Freeform message text is not used as a policy source. |
| Fixed untrusted cases | The designated untrusted example and specific instruction patterns withhold a draft. This is a finite rule check, not a general injection defense. |
| Review and export | Export requires review. Input changes, draft edits, new case or language results, and reset revoke prior approval. Blocked, stale, empty, or overlong drafts cannot be approved or exported. |
The full test file below contains the individual assertions and values. It tests local rules and state transitions; it does not evaluate the meaning of arbitrary customer questions.
Observed in the browser.
- Changing a policy changes the result. An edited shipping value appeared in both the draft and its displayed evidence after rebuilding.
- Missing information withholds the draft. The unsupported example showed a hold state instead of an invented answer.
- A conflict remains visible. Return windows of 7 and 14 days blocked drafting and exposed the conflicting policy values.
- Resolving the conflict updates the Korean reply. Changing both return windows to 14 days restored the draft, with 14 days reflected in Korean.
- Review controls the text export. A draft required review before its local TXT export. The downloaded result contained the reviewed draft and supporting evidence.
- An edit revokes approval. Editing a reviewed draft required a new review before export became available again.
These observations establish that selected controls affect the underlying draft and review state. They do not establish speed improvements, translation quality across unseen text, or readiness for customer records.
Run the same rule checks.
Download these three files into one folder. The published test copy uses relative imports; its test cases and assertions match the project suite. A Node.js runtime with the built-in test runner is required.
- prototype-tests.mjs: the 17 existing tests.
- engine.mjs: policy rules and review state.
- sample-data.mjs: fictional defaults and cases.
node --test prototype-tests.mjsThe original command from the project workspace:
node --test work/daewon-tests/engine.test.mjsRe-running the suite checks the files you downloaded. The recorded result above is dated; future code changes should be followed by a new run. Browser interactions and download contents require separate browser checks.
Live-model evaluation is still ahead.
The next proposed implementation adds approved-document retrieval and server-side Claude drafting. That work needs a separate evaluation of factual claim support, policy constraints, citation support, correct escalation, bilingual quality, latency, and cost.
A consented merchant pilot is also planned, not completed. The product brief describes the evaluation approach and a target of three shops over two weeks. No model-performance or customer-outcome results are reported here.