FEATUREDLatent Space· rssEN23:37 · 08·21
→Simulation as the New Scaling Law — Joon Sung Park, Simile AI
Joon Sung Park, co-founder and CEO of Simile AI, walked through the company's roadmap on Latent Space. Simile just raised a $200M Series B at a $2B valuation led by GreenOaks and Index Ventures, with backers including Fei-Fei Li and Andrej Karpathy. Their core product trains behavioral foundation models on long-form interviews, transaction data, and randomized controlled trials to build digital twins of real people. In a 1,000-person study, the models reproduced human behavior and attitudes with 85% accuracy—comparable to how consistently the same person retakes a test. Joon argues frontier models are too rational to simulate real humans, who make mistakes and hold biases. Simile is already running tens of millions of simulations for Fortune 100 clients like CVS, replacing expensive human focus groups. The long-term ambition is to simulate all 8 billion people to test products, policies, and even study climate change or UBI. The post does not disclose the specific evaluation benchmark behind the 85% figure.
#Simile AI#Joon Sung Park#GreenOaks
why featured
Featured · importance 82 · hook + knowledge + resonance
editor take
Simile trains behavioral models on interviews and transaction data, hitting 85% accuracy in a 1,000-person study, but the eval benchmark isn't disclosed.
sharp
This caught my eye because Simile just raised $200M at a $2B valuation, with Fei-Fei Li and Andrej Karpathy on the cap table. The pitch is straightforward: build digital twins of real people to replace expensive human focus groups. CVS is already a customer.
Joon Sung Park's core argument is that frontier models are too rational to simulate humans who make mistakes and hold biases. So instead of prompting GPT or Claude, Simile trains behavioral foundation models on long-form interviews, transaction data, and randomized controlled trials. In a 1,000-person study, their models reproduced human behavior and attitudes with 85% accuracy—comparable to how consistently the same person retakes a test.
I'd discount this a bit. The post doesn't disclose the evaluation benchmark behind that 85% figure, so we don't know what tasks were tested or what baseline they're comparing against. Privacy and consent around interview and transaction data also go unaddressed. Joon's long-term vision—simulating all 8 billion people to test policies, study climate change, or model UBI—reads more like ambition than roadmap.
The near-term value is real: for consumer goods companies running large-scale user research, this is cheaper than traditional focus groups. Just don't read it as "digital twins are accurate enough to replace human decision-making." Right now it looks more like a narrow survey-replacement tool.
HKR breakdown
hook ✓knowledge ✓resonance ✓