TomekKaszyński
Figuring out what intelligence actually requires. Not by scaling LLMs. Somewhere between biology and brute-force engineering.
On ARC-AGI-3, the benchmark every frontier LLM scored under 1% on at launch: GPT-6 Astra inside my harness clears a game 8/8 cold at 7,704 output tokens. Inside my loop before the slicing it needed 35,842. ARC Prize's verified run of the same model costs $18,817.
No LLM does the thinking. The harness does; the model only picks, and every decision leaves a receipt. An open model a fraction of the size makes the same picks.
Also co-founding a fintech.