4 ms·
I think the thing I'm most excited about is the increase in _user prompting_. If I give a poorly constrained/ambiguous prompt, I don't want the model one-shott
by jumploops 15d ago
I think the thing I'm most excited about is the increase in _user prompting_.
If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right.
The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.
It's a tough balance to get right, and although this has been possible to achieve with additional prompting on existing models, I find that the agents often lean too hard into the "ask questions" mode.
Hopefully this model has the right balance, or at least better?
- weird-eye-issue 15d agoFable does a great job from my terrible prompts when coding
- dannyw 15d agoAnecdotal experiences from my external early testing of Astra: if you love Sol (like I do) and wished it was smarter at everything, but especially better at high-level tasks and discussions; I think you'll LOVE Astra. Astra retains the best parts and overall 'grounded collaborator and executor' of Sol in my testing (harness: codex CLI); while being a significant leap in capabilities & higher-level thinking. When you prompt it like a technical collaborator, I've found Astra to be extremely consistent in staying as a collaborator, and not being over-eager, over-achieving or doing work that you haven't asked it to. When you ask it to one-shot something, or explicitly ask it to make decisions, it will of course make its own assumptions and decisions, and generally very well. Astra is also excellent at instruction following and respecting the guidance and steers boundaries you have. ^OpenAI does not review, limit, or tell me what to say; opinions are my own experiences.
- NorthSouthNorth 14d agoThis is quite exciting. Sol for me has been the absolute best model yet. I find myself using it 95% of the time even though I have access to Fable. Is the speed the same as 5.6 sol?
- fps-hero 14d agoFellow excited. Sol has been revolutionary for me. It’s been the first time I’ve let a model “go ham” on a production code base and it actually worked, and also didn’t bankrupt my company in the process. I’m far more excited for the trajectory that OpenAI has chosen. I’ve been listening to mates whose companies adopted Claude wholesale, only for their AI use to become a double digit percentage of their salary. I get that AI is a force multiplier, but that level of expense isn’t a path to mass adoption. Honestly, this feels like a real revolution now in a way that 90s kid never really experienced. We grew up with technological progression, we never experienced the obsolescence of skill. The internet revolution made for more skills and innovation, it didn’t obsolete entire careers. My kids are almost certainly going to grow up knowing less but being capable of more. Imagine being a 1950s “human calculator” on the dawn of a computer revolution. That’s what it feels like right now.
- chrisweekly 14d ago> "Imagine being a 1950s “human calculator” on the dawn of a computer revolution. That’s what it feels like right now." My personal / family history is a real-world example of that evolution. My grandfather was a "computer", my father was a traditional "programmer" (lots of Perl), and I'm a SWE / frontend architect / budding "AI Engineer".
- natsucks 14d ago> I’ve been listening to mates whose companies adopted Claude wholesale, only for their AI use to become a double digit percentage of their salary. sigh...yep. that's us.
- wrsh07 14d agoOff topic, but ooc what do you do such that you get early access to the models?
- nullbio 15d agoThis is spot on. A collaborator is exactly what real AGI is. It will figure out the perfect questions to ask, in the perfect order, by intelligently assessing the entire solution and problem space upfront, so when you leave it to go off on its own it isn't making stupid decisions for you. They really need to make this work in Codex. Claude Code has had a multi-select refinement tool since forever.
- apsurd 15d agoonly if you specifically ask Claude to use AskUserQuestion tool. otherwise i would hardly say claude acts as a collaborator naturally.
- weird-eye-issue 14d agoFor relatively complex new features it will automatically do this
- embedding-shape 15d ago> They really need to make this work in Codex. Claude Code has had a multi-select refinement tool since forever. I think this already exists in Codex? If you use "/plan" and something is unclear or ambiguous, Codex will ask you and present choices, and let you enter your own custom answer. Then it'll iterate like this until the plan is clear and ambiguous. Isn't this what you're talking about? If so, it has existed for a long time in Codex. Overall I agree with you though, all the models currently don't have the right hunches nor the right approach about when things are clear enough or not.
- lmf4lol 14d agoyes. There is plan mode in codex. I use it all the time. It will come up with a set of questions to clarify things
- enraged_camel 15d ago>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever. I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it. I watched some reviews today and came way with the impression that Astra is not better than Sol in this regard. You still have to be very specific with your instructions. For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed. But it won't commit unless you explicitly say so. That sounds like a very tedious way of working with AI agents, but I understand some people want a high level of control.
- StevenWaterman 15d ago> For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed. That's exactly what i want to happen. I hate when it assumes my direct question was an indirect instruction
- embedding-shape 15d agoThis is one of those things that won't ever be "solved" as people just want different things here, hence we can steer the models with the system prompt. I'm mostly the same as you, I don't want the model to assume things, or act on implicit "directions". But then also, sometimes I do, and I myself might not always know when what approach is best.
- amazingman 13d agoIndeed some people think they want a machine guessing at your intentions and acting upon that guess. Those people Are wrong in at least 2 directions: that it is what they want, and their implicit assumption that it could possibly be safe.
- diroussel 14d agoIndeed, I don’t ask rhetorical questions to an AI. They are of doubtful use when talking verball ly to a human, less good in online discussions, and totally unnecessary for agents.
- firemelt 15d agothis before executing I ask my ai to discuss what I mean/intention
- speleding 14d agoThat balance probably depends on the human, and the context. If you are a beginner in a field the model should not assume you know what you are doing. On the other hand, for an expert it should try to work out what you mean with your vaguely worded order. What I think should happen is that it should update its memory with notes on the proficiency level of the user, so it gets the balance right over time. This is a problem if you allow your kids to use your ChatGPT account for homework (and silly pictures), like I do.
- Zambyte 14d ago> What I think should happen is that it should update its memory with notes on the proficiency level of the user, so it gets the balance right over time. Which can also involve just asking for the users level of experience
- cerol 14d agoand that's exactly what most people don't want. they want some ultra-intelligent being that can do marvelous things, and they can claim the credit on it