OMEGA-Reasoner (POC)

⚠️ This is a proof-of-concept β€” a small-scale demonstration of a self-contained LLM-driven agent framework. It is not a frontier model and is not intended to compete with production models on public leaderboards such as DeepSWE.

Honest positioning

Property This repo DeepSeek-V4.1-Flash (Top-1 DeepSWE)
Parameters β‰ˆ 100K (configurable) 35B-class
Training data none (synthetic / hand-coded) trillions of tokens
Pass@1 on DeepSWE not measured 74.2
Status research POC production model

We do not submit to the official DeepSWE leaderboard. The integration code (deep_swe.py) is provided for educational and self-evaluation purposes only.

What this model is

OMEGA-Reasoner is a self-contained language world model + autonomous agent framework. It implements:

  • HMoE-D reasoner: hierarchical mixture-of-experts with dynamic routing
  • AGI_AgentWorld: full cognitive agent with episodic / semantic / procedural memory, ReAct loop, Tree-of-Thoughts planning, Reflexion-style self-correction, and constitutional guardrails
  • AgentWorldBench adapter: 5-dim scoring (format, factuality, consistency, realism, quality)
  • DeepSWE adapter: Harbor-format task loader + verifier runner

Why this repo exists

To show that a complete agent framework β€” including memory, planning, tool-use, self-correction, and benchmark adapters β€” can be implemented from scratch without any external LLM API. It is meant as a research artifact, not a competitive submission.

How to use

from evo.agi_agent_world   import AGIAgentWorld
from evo.agentworld_bench  import load_all, WorldModelAgent, evaluate
from evo.deep_swe          import load_all_tasks, run_benchmark

# 1) Run a single goal
agi = AGIAgentWorld()
result = agi.run("Buy coffee by noon")

# 2) Evaluate on AgentWorldBench (synthetic data)
records = load_all("data/agentworldbench_synth")
rep     = evaluate(WorldModelAgent(), records)
print(rep.overall)

# 3) Run on DeepSWE tasks (synthetic seed data)
tasks   = load_all_tasks("data/deepswe_synth")
summary = run_benchmark(tasks)
print(summary.pass_at_1)

Files

  • agi_agent_world.py β€” autonomous agent (memory, planning, tools)
  • omega_arch.py β€” HMoE-D model definition
  • omega_nlu.py β€” NLU engine (sentences, frames, FOL)
  • omega_math.py β€” math reasoning primitives
  • omega_physics.py β€” physics reasoner
  • omega_meta.py β€” meta-cognition / audit
  • omega_quant.py β€” quantization utilities
  • agentworld_bench.py β€” AgentWorldBench loader + 5-dim scorer
  • deep_swe.py β€” DeepSWE loader + verifier
  • tests/ β€” unit tests (78+ tests, all passing)

Test results

The repository ships with 78+ unit tests covering metrics, agents, benchmarks, and the cognitive core. All tests pass on the included synthetic fixtures.

Limitations

  • The underlying HMoE-D model is a demonstration with no language pretraining. Its predict() is template-based, not learned.
  • Pass@1 on DeepSWE is not claimed β€” the agent is heuristic.
  • Numerical / symbolic reasoning is rule-based, not learned.

Citation

If you use this code in research, please cite the framework only (no leaderboard claim):

@software{omega_reasoner_poc,
  title  = {OMEGA-Reasoner: a self-contained LLM-agent framework (POC)},
  year   = {2026},
  url    = {https://huggingface.co/<your-org>/omega-reasoner-poc},
}

License

Apache-2.0.

Downloads last month
22
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Datasets used to train bbkdevops/omega-reasoner-poc