Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
lout332
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
I Failed 3 Times Building This with AI. In 2026, It Took Days
(luisfernandoyt.makestudio.app)
3 points
by
lout332
7mo ago
|
0 comments
2.
▲
by
lout332
7mo ago
lets meet https://luisfernandoyt.makestudio.app/
3.
▲
Conversations with an AI That Argues Back
(luisfernandoyt.makestudio.app)
1 points
by
lout332
7mo ago
|
0 comments
4.
▲
Conversations with AI: What I Learned About Myself
(luisfernandoyt.makestudio.app)
1 points
by
lout332
7mo ago
|
0 comments
5.
▲
by
lout332
7mo ago
agreed. local is key. what's your setup?
6.
▲
I gave an AI access to my psychology. Documenting the experiment
(github.com)
3 points
by
lout332
7mo ago
|
3 comments
7.
▲
by
lout332
7mo ago
I've been running an experiment for the past few months: full AI integration into my daily life. Not a chatbot I use occasionally. A "symbiotic agent" that reads two files at every session: one with my identity, psychology, a
8.
▲
by
lout332
8mo ago
Thanks, noted. Will fix.
9.
▲
by
lout332
8mo ago
You're right about the state sync issues with some models. The lighter models (especially Llama) struggle with tracking game state. I've added more Gemini options which handle this better. The research data used controlled AI-vs-A
10.
▲
by
lout332
8mo ago
Full game logs are in data_public/comparison/ on GitHub. Each JSON has the complete game state, moves, and messages across all 162 games. https://github.com/lout33/so-long-sucker
11.
▲
by
lout332
8mo ago
the interactive demo uses lighter models for cost reasons. The research data (162 games, 90% Gemini win rate) came from longer AI-vs-AI games where strategic depth emerged over 50+ turns. Short games with a human tend to expose the models&#
12.
▲
by
lout332
8mo ago
Sure, no problem, I added a new section explaining the game
13.
▲
by
lout332
8mo ago
Fixed - donation flow no longer blocks the game. Thanks for the report.
14.
▲
by
lout332
8mo ago
Game logs are in data_public/comparison/ - each JSON has the full game state, moves, and messages. For example, check gemini_vs_all_7chips.json to see the alliance bank betrayals in action.
15.
▲
by
lout332
8mo ago
Full code and raw data: https://github.com/lout33/so-long-sucker
16.
▲
by
lout332
8mo ago
Not yet, but I'd be interested in collaborating on one. The dataset (162 games, 15K+ decisions, full message logs) is available. If you know anyone in AI Safety research who'd want to co-author, I'm open to it.
17.
▲
by
lout332
8mo ago
Fair point. The core simulation and data collection was done programmatically - 162 games, raw logs, win rates. The analysis of gaslighting phrases and patterns was human-reviewed. I used LLMs to help with the landing page copy, which I sho
18.
▲
by
lout332
8mo ago
Used Kimi K2 (the main reasoning model). For the thinking space - we gave all models access to a think tool they could optionally call for private reasoning. Gemini used it heavily (planning betrayals), GPT-OSS never called it once. The int
19.
▲
by
lout332
8mo ago
> "Thanks for trying it! I'll look into the 'Pile not found' error and fix it. > > For rules, here's a 15-min video tutorial: https://www.youtube.com/watch?v=DLDzweHxEHg > > On autoro
20.
▲
Which AI Lies Best? A game theory classic designed by John Nash
(so-long-sucker.vercel.app)
195 points
by
lout332
8mo ago
|
80 comments
21.
▲
by
lout332
8mo ago
We used "So Long Sucker" (1950), a 4-player negotiation/betrayal game designed by John Nash and others, as a deception benchmark for modern LLMs. The game has a brutal property: you need allies to survive, but only one player
22.
▲
I built an interactive simulator to explore AI futures (2025-2030)
(ai-futures.vercel.app)
1 points
by
lout332
9mo ago
|
0 comments
23.
▲
Show HN: Claude Life Assistant – AI accountability partner for Claude Code
(github.com)
3 points
by
lout332
9mo ago
|
0 comments
24.
▲
Claude Life Assistant: Personal accountability coach in your filesystem
(github.com)
1 points
by
lout332
9mo ago
|
0 comments
25.
▲
Show HN: Canvas and Agent: Cursor and Canvas makes a baby
(canvas-agent.vercel.app)
1 points
by
lout332
10mo ago
|
0 comments
26.
▲
LLM council web ready to use version
(ai-brainstorm-blue.vercel.app)
1 points
by
lout332
10mo ago
|
0 comments
27.
▲
Party in the AI Lab (Parody of Parody "Party in the CIA." By Weird Al Yankovic) [video]
(youtube.com)
1 points
by
lout332
10mo ago
|
0 comments
28.
▲
Party in the AI Lab – AI Safety Parody (Weird Al Style) [video]
(youtube.com)
1 points
by
lout332
11mo ago
|
1 comments
29.
▲
by
lout332
11mo ago
A sharp and hilarious parody that tackles AI research culture, alignment debates, and safety concerns through comedy. This Weird Al-style musical parody resonates with the current state of AI development, poking fun at researcher competitio
30.
▲
Data Viz: Mapping Model Performance on Reasoning vs. Honesty Benchmarks
(claude.ai)
1 points
by
lout332
1y ago
|
1 comments
More ›