TomekKaszyński

Now

Figuring out what intelligence actually requires. Not by scaling LLMs. Somewhere between biology and brute-force engineering.

On ARC-AGI-3, the benchmark every frontier LLM scored under 1% on at launch: GPT-6 Astra inside my harness clears a game 8/8 cold at 7,704 output tokens. Inside my loop before the slicing it needed 35,842. ARC Prize's verified run of the same model costs $18,817.

No LLM does the thinking. The harness does; the model only picks, and every decision leaves a receipt. An open model a fraction of the size makes the same picks.

Also co-founding a fintech.

Projects
  1. Tractatus

    no llm does the thinking
Elsewhere