Skip to content

Benchmark

No model enters OYYO on reputation.

The OYYO Benchmark is being built as a reproducible way to measure whether a model can actually do the work, on real hardware, with real constraints.

In developmentยท 0.1 onwards

Capability families

  • General reasoning
  • Business execution
  • Coding and terminal work
  • Tools, functions and agents
  • Documents, spreadsheets, presentations
  • Language and semantic translation
  • Memory and temporal knowledge
  • Image, audio and video understanding
  • Growth and marketing execution
  • Security and reliability
  • Hardware efficiency

Metrics

Quality and cost, measured together.

A model that is excellent but unusable on the hardware you own is not a solution.

  • Task success
  • Quality
  • Repeatability
  • Latency
  • Time to first token
  • Tokens per second
  • RAM
  • VRAM
  • Load time
  • Context
  • Crash and fallback behaviour
  • Energy where measurable

Benchmark results will be published as the harness matures. Until then, no OYYO page claims a score.

Think. Create. Execute.

Evidence before branding.

OYYO is being built in public against a versioned release line. Join the early access list and follow the roadmap.