Spaces:
Configuration error
Configuration error
| title: README | |
| emoji: π¦ | |
| colorFrom: gray | |
| colorTo: blue | |
| sdk: static | |
| pinned: false | |
| # Dissei Data | |
| **Outcome-graded RL environments for institutional financial judgment.** | |
| Dissei builds reinforcement-learning environments from real institutional credit | |
| decisions β each one a documented transaction rebuilt as an agentic task, worked with | |
| analyst tools and graded against what actually happened. | |
| - **Point-in-time by construction.** A model sees only what was knowable at the decision | |
| date; every exhibit is anchor-gated, so hindsight is excluded structurally β not by an | |
| instruction it can ignore. | |
| - **Graded, not checked.** Reward decomposes into fail-closed critical gates, an | |
| expert-authored graded rubric, a contradiction penalty, and a retrieval modulator that | |
| rewards grounding in the record. | |
| - **Grading automation, disclosed.** Rubrics are authored and anchored by credit | |
| practitioners who carried real risk; execution runs on LLM judges at temperature 0 that | |
| apply the sealed rubric and never invent criteria. Answers are graded twice β against | |
| best practice, and against the realized outcome. | |
| - **Sealed by architecture.** Answer material is time-locked and anchor-gated; every asset | |
| ships with a per-asset contamination statement and right-to-train lineage. | |
| - **Operators, not an annotation farm.** The corpus is the by-product of an operating | |
| lending business β incentive-bearing decisions with real consequences, not opinions | |
| commissioned for a dataset. | |
| Environments (RL post-training) and sealed evaluations are available to frontier labs | |
| **under agreement**. We do not publish tasks, rubrics, graders, or leaderboards. | |
| **Methods & findings:** https://dissei.ai/research | |
| **Contact:** info@dissei.credit Β· Mutual NDA before anything sensitive. | |