README / README.md
eb-dissei's picture
Add space metadata front-matter
4b258eb verified
|
Raw
History Blame Contribute Delete
1.8 kB
metadata
title: README
emoji: 🏦
colorFrom: gray
colorTo: blue
sdk: static
pinned: false

Dissei Data

Outcome-graded RL environments for institutional financial judgment.

Dissei builds reinforcement-learning environments from real institutional credit decisions — each one a documented transaction rebuilt as an agentic task, worked with analyst tools and graded against what actually happened.

  • Point-in-time by construction. A model sees only what was knowable at the decision date; every exhibit is anchor-gated, so hindsight is excluded structurally — not by an instruction it can ignore.
  • Graded, not checked. Reward decomposes into fail-closed critical gates, an expert-authored graded rubric, a contradiction penalty, and a retrieval modulator that rewards grounding in the record.
  • Grading automation, disclosed. Rubrics are authored and anchored by credit practitioners who carried real risk; execution runs on LLM judges at temperature 0 that apply the sealed rubric and never invent criteria. Answers are graded twice — against best practice, and against the realized outcome.
  • Sealed by architecture. Answer material is time-locked and anchor-gated; every asset ships with a per-asset contamination statement and right-to-train lineage.
  • Operators, not an annotation farm. The corpus is the by-product of an operating lending business — incentive-bearing decisions with real consequences, not opinions commissioned for a dataset.

Environments (RL post-training) and sealed evaluations are available to frontier labs under agreement. We do not publish tasks, rubrics, graders, or leaderboards.

Methods & findings: https://dissei.ai/research Contact: info@dissei.credit · Mutual NDA before anything sensitive.