Jev tested on 8,054 NASA Kepler signals: 54.2% accuracy, loses to a simple 3-rule baseline
This_Cell_1829 · reddit · 2026-09-20
Reddit user ThisCell1829 ran a retrospective classification test of Jev on 8,054 historical Kepler Objects of Interest, giving it 21 measurements per signal to classify each as confirmed planet, false positive, or candidate—saving all predictions before checking NASA Exoplanet Archive labels.
Results: Jev scored 54.2% overall, below a simple 3-rule baseline (64.4%) and only above always-guessing-false-positive (49.0%). It was strong at catching false positives (89.6%), but extremely conservative on confirmed planets: out of 2,731, it used the "confirmed planet" label only 8 times—all 8 correct—missing the other 2,723, mostly calling them candidates. It behaved more like a cautious false-positive filter than a general classifier.
Other numbers: 8,054 calls, zero failures, 338ms median latency, roughly $0.36 total cost. The author also made a short visualization using the real Kepler field.
More from Models
- Grok Voice Transcribe 2.0 cuts short-phrase WER from 20.6% to 6.8%, tops streaming STT — nima_owji · 2026-09-20
- FT: AI chatbots give wrong answers to financial queries 'most of the time' — SatelliteNetSec · 2026-09-20
- Testing Jev on counting letters: it returns probabilities and still flubs obvious questions — eptwts · 2026-09-20
- Stab-a-figure test: Fable refuses all 20 trials while Astra complies in 85% — AaronBergman18 · 2026-09-20
- DiffusionGemma 26B-A4B turned into a local System One fast-decision model via vLLM patch — solyarisoftware · 2026-09-20
- Jev beats GPT-5.6 at predicting book tastes, 53x cheaper and 25x faster — venturetwins · 2026-09-20