Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
philipbjorge
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
philipbjorge
22d ago
That's how it started for us. Make sure you achieve and maintain great profitability, PE operates very differently from VC.
2.
▲
by
philipbjorge
24d ago
> When you actually put them to the test you see 500k tokens of reasoning with "Actually..." and "Wait..." in every third paragraph of their reasoning trace. I've wondered if this is part of why we don't see
3.
▲
by
philipbjorge
27d ago
I'm writing much more secure software and infrastructure (e.g. IAM now gets configured well).
4.
▲
by
philipbjorge
2mo ago
Wow, I would love to read an interview series based on this!!
5.
▲
by
philipbjorge
3mo ago
It’s been 25 years since I was in school and this was my experience. Unsure if it’s changed…
6.
▲
by
philipbjorge
3mo ago
I’ll second this, I had to send this message… > based on our discussions so far, I think we’d only need... Which they appear to have but I can't dig in on mobile.
7.
▲
by
philipbjorge
3mo ago
I'll go further and note that some of the optimizations I've seen in rtk for things like `git status` have actually bubbled up into the model layer -- Codex is regularly making tool calls like `git status --short` instead of `git
8.
▲
by
philipbjorge
4mo ago
I can't find the relevant issues in their repo, but I've been somewhat skeptical of their tool over-reporting token savings and there are many issues to that effect in the repo. I'm not likely to install it again in my latest
9.
▲
by
philipbjorge
4mo ago
I've done some pretty incredible things with LLMs. If this were sqlite with its exhaustive test suite... OK, I can see it. It's hard for me to see this not becoming a pile of slop, but hey, maybe I'm wrong
10.
▲
by
philipbjorge
4mo ago
Been loving pi and codex lately. Good to build resiliency and self sufficiency into these systems.
11.
▲
by
philipbjorge
4mo ago
Can you fill me in on how this impacts conductor? How are they using `claude -p`?
12.
▲
by
philipbjorge
4mo ago
Tracking with `ccusage`, I pretty easily hit $2000/mo in API equivalent credits and while I'd consider myself a power user, I'm a responsible one that's generally always in the loop. If I were using `claude -p`, this wou
13.
▲
by
philipbjorge
4mo ago
This seems like less of a today thing and more of an ancient human tendency. A lot of Buddhist practice is basically trying to train against immediately collapsing reality into self/other, right/wrong, craving/aversion. Pract
14.
▲
by
philipbjorge
5mo ago
I haven't really shared what I use, I'm still deciding if that's something I want to do. To get an idea of what I'm talking about, you could install https://github.com/obra/superpowers/ into bo
15.
▲
by
philipbjorge
5mo ago
Ahh good point -- I've handled this by switching my harness to `pi` but recognize that may not be for everyone and doesn't directly address OP's question.
16.
▲
by
philipbjorge
5mo ago
What I found was that I *strongly* preferred Claude Code with its defaults. Codex was almost unusable to me -- It would spit out a 4-5 page plan where it kept repeating itself, where Claude would give me a crisp 1-2 pager I could actually r
17.
▲
by
philipbjorge
5mo ago
You might search for a concept like `/handoff` that's in ampcode. I'm sure someone's built a skill for just this.
18.
▲
by
philipbjorge
5mo ago
So happy to have diversified my model providers this past couple of weeks. GPT-5.5 has had no trouble slotting into Opus workloads. Will be fun to try out more of the models as time goes on to build some resiliency into my engineering workf
19.
▲
by
philipbjorge
5mo ago
I’ve been comparing Claude Code and Codex extensively side by side over the past couple of weeks with my favorite prompting framework superpowers… From my perspective, Claude Code is decidedly not better than Codex. They’re slightly differe
20.
▲
by
philipbjorge
5mo ago
gpt 5.4 has been performing great in my harness.
21.
▲
by
philipbjorge
8mo ago
This looks remarkably similar to https://github.com/vercel-labs/agent-browser How is it different?
22.
▲
by
philipbjorge
10mo ago
Until it's happened to you, it sounds unbelievable Sorry about all the broken plastic on the trim -- That's also very familiar...
23.
▲
by
philipbjorge
11mo ago
> We asked for a bill with the standard CPT codes. No reply. Asked again. “Oh, we meant to send it. We upgraded our computers five months ago and nothing works.” Uh-huh. Finally got the CPT codes. I work in healthcare RCM. I have no trou
24.
▲
by
philipbjorge
11mo ago
We had a similar realization here at Thoughtful and pivoted towards code generation approaches as well. I know the authors of Skyvern are around here sometimes -- How do you think about code generation with vision based approaches to agent
25.
▲
by
philipbjorge
1y ago
Important is subjective — In the healthcare space, I’d make the claim that most applications don’t expose themselves correctly (native or web). CV and direct mouse/kb interactions are the “base” interface, so if you solve this problem,
26.
▲
by
philipbjorge
1y ago
If you're looking to test an LLMs ability to solve a coding task without prior knowledge of the task at hand, I don't think their benchmark is super useful. If you care about understanding relative performance between models for s
27.
▲
by
philipbjorge
1y ago
devcontainers extension was a year out of date up until the last month or something? sorry, this is from memory, but definitely not 100% compatibility.
28.
▲
by
philipbjorge
2y ago
I haven't followed this closely, but I assumed it was related to a foreign entity having the ability to hyper-target content towards said 17 year olds (and the entire userbase in general) -- A modern form of psychological warfare.
29.
▲
by
philipbjorge
2y ago
The latest research I’ve pulled suggests that DEXA scans are fairly inaccurate and aren’t a reliable way to measure body composition even for the same person across time. MRI is the gold standard, everything else is pretty loosely goosey. S
30.
▲
by
philipbjorge
2y ago
They report that they trained the model to count pixels and based on accurate mouse clicks coming out of it, it seems to be the case for at least some code path. > When a developer tasks Claude with using a piece of computer software and
More ›