Spaces:
Configuration error
Configuration error
metadata
title: README
emoji: 🏦
colorFrom: gray
colorTo: blue
sdk: static
pinned: false
Dissei Data
Outcome-graded RL environments for institutional financial judgment.
Dissei builds reinforcement-learning environments from real institutional credit decisions — each one a documented transaction rebuilt as an agentic task, worked with analyst tools and graded against what actually happened.
- Point-in-time by construction. A model sees only what was knowable at the decision date; every exhibit is anchor-gated, so hindsight is excluded structurally — not by an instruction it can ignore.
- Graded, not checked. Reward decomposes into fail-closed critical gates, an expert-authored graded rubric, a contradiction penalty, and a retrieval modulator that rewards grounding in the record.
- Grading automation, disclosed. Rubrics are authored and anchored by credit practitioners who carried real risk; execution runs on LLM judges at temperature 0 that apply the sealed rubric and never invent criteria. Answers are graded twice — against best practice, and against the realized outcome.
- Sealed by architecture. Answer material is time-locked and anchor-gated; every asset ships with a per-asset contamination statement and right-to-train lineage.
- Operators, not an annotation farm. The corpus is the by-product of an operating lending business — incentive-bearing decisions with real consequences, not opinions commissioned for a dataset.
Environments (RL post-training) and sealed evaluations are available to frontier labs under agreement. We do not publish tasks, rubrics, graders, or leaderboards.
Methods & findings: https://dissei.ai/research Contact: info@dissei.credit · Mutual NDA before anything sensitive.