--- title: README emoji: ๐Ÿฆ colorFrom: gray colorTo: blue sdk: static pinned: false --- # Dissei Data **Outcome-graded RL environments for institutional financial judgment.** Dissei builds reinforcement-learning environments from real institutional credit decisions โ€” each one a documented transaction rebuilt as an agentic task, worked with analyst tools and graded against what actually happened. - **Point-in-time by construction.** A model sees only what was knowable at the decision date; every exhibit is anchor-gated, so hindsight is excluded structurally โ€” not by an instruction it can ignore. - **Graded, not checked.** Reward decomposes into fail-closed critical gates, an expert-authored graded rubric, a contradiction penalty, and a retrieval modulator that rewards grounding in the record. - **Grading automation, disclosed.** Rubrics are authored and anchored by credit practitioners who carried real risk; execution runs on LLM judges at temperature 0 that apply the sealed rubric and never invent criteria. Answers are graded twice โ€” against best practice, and against the realized outcome. - **Sealed by architecture.** Answer material is time-locked and anchor-gated; every asset ships with a per-asset contamination statement and right-to-train lineage. - **Operators, not an annotation farm.** The corpus is the by-product of an operating lending business โ€” incentive-bearing decisions with real consequences, not opinions commissioned for a dataset. Environments (RL post-training) and sealed evaluations are available to frontier labs **under agreement**. We do not publish tasks, rubrics, graders, or leaderboards. **Methods & findings:** https://dissei.ai/research **Contact:** info@dissei.credit ยท Mutual NDA before anything sensitive.