12 ms·
I Spent a Week with Gemini Pro 1.5–It's Fantastic
- Solvency 3y agoHow can Google so thoroughly embarrass themselves on the image front and then do well on text?
- itchy_spider44 3y agoBecause engineers make great things and management tells them to make it woke
- hersko 3y agoThey are not doing well on text.... https://deepnewz.com/politics/google-s-woke-gemini-ai-slammed-bias-amplifying-hamas https://deepnewz.com/politics/google-s-woke-gemini-ai-slamme...
- deleted 3y ago[deleted]
- wredue 3y ago[flagged]
- Solvency 3y ago...what?
- wredue 3y ago[flagged]
- artninja1988 3y agoWeird strawman about what actually happened. If you prompted "ginger women" or even the names of two Google confounders you got a completely different race for no reason. Embarrassing
- platelminto 3y agoGPT-4 Turbo has a context window of 128k tokens, not 32k as the article says.
- dshipper 3y agoFixed!! Thanks for finding that
- inciampati 3y agoThe sibling comment explained. The web client is 16k or 32k, but you can pay their API $1.28 per interaction and use a full context on GPT-4-turbo.
- Workaccount2 3y agoI believe with the API you can get 128k, but using the app or web client it is 32k. This might have changed though.
- criddell 3y agoI kind of love the idea of feeding the text of entire books to an AI. More often than I’d like to admit, I’ll be reading a novel and find myself not remembering who some character is. I’d love to be able to highlight a name in my ereader and it would see that I’m 85 pages into Neuromancer and give me an answer based on that (ie no spoilers). Or have a textbook that I can get some help and hints while working through problems and get stuck like you might get with a good study partner.
- TrueGeek 3y agoI'm the opposite. In movies (and real life to an extent) I have trouble telling people apart. I'd love a closed caption like service that just put name tags over everyone.
- golergka 3y agoWhen I first watched Departed, I didn't realise that Matt Damon and Leonardo DiCaprio's characters are different people until the third act. It was very confusing.
- neltnerb 3y agoHah, I thought I was the only one. I'm not particularly face blind either... something about the era I guess.
- antifa 3y agoLoudest thing online is people complaining about minorities existing in TV shows. A much bigger, real problem, is when they regularly cast 3 characters with almost exactly the same skin tone, skin texture, hair color, hair length, hair style, eye color, clothing style, face shape, voice, body type. Then the character's name is given once (if given) and you see several scenes without the character. Several scenes with the clones. Sudden scene with 2 clones in the same scene. You really couldn't give these three white people a different hair style? Make one wear glasses?
- ryandrake 3y ago
- Sakos 3y agoI love the potential of having such a big context window, but I'm concerned about who will get access to it (or rather who won't get access to it) and what it will cost or who will pay for it.
- criddell 3y agoYou could have asked the same question in 1964 when IBM released the System/360. Nice computer, but who will pay for it and who will have access to it? I think it’s inevitable that these AI’s will end up costing almost nothing to use and will be available to everybody. GPT-4 is already less than $1 / day and that’s only going to go down.
- Sakos 3y agoI didn't ask it then. I wasn't even alive. I'm asking now. It is a legitimate concern and continues to be a concern with these AI models and you whining that "oh but 50 years ago bla bla" (as if these two things are in any way comparable) doesn't change anything about the risks of selective access.
- eesmith 3y agoI do not think that's a good example. A lot of people jumped on the chance to buy a 360. Here are some quotes from https://archive.org/details/datamation0064unse_no5/page/68/mode/2up?q=%22System+360%22 https://archive.org/details/datamation0064unse_no5/page/68/m... : > The internal IBM reaction could be characterized as quiet, smug elation. One office is supposed to have sold its yearly quota on A (Announcement) -Day. In Texas, a man allegedly interrupted the 360 presentation to demand he be allowed to order one right then . . . which sounds like a combination plant and a new version of the rich Texan jokes. ... > the 360 announcement has to worry the competition considerably . . . partly because anything new from IBM creates an automatic bandwagon effect, partly because the completeness of the new line offers less reason for people to look outside. ... > another feels that the economic incentive (rental cuts of 50 per cent for 7080, 7090) will force him down the 360 route. And he thinks 360 interrupt features will open the door to real-time applications which can be approached on an incremental basis impossible before. ... > One maverick doesn’t share the enthusiasm of his company, which ordered “plenty” of 360’s within an hour of the announcement, without price agreements. And from https://archive.org/details/bitsavers_Electronic9640504_98093295/page/34/mode/2up?q=%22System+360%22 https://archive.org/details/bitsavers_Electronic9640504_9809... : > Other computer manufacturers profess to be unworried, but International Business Machines Corp. has received millions of dollars in orders for its system/360 computer [Electronics, April 20, p. 101]. Here are some of the people who paid for it in 1964: https://archive.org/details/sim_computers-and-people_1964_13_index/page/30/mode/2up?q=%22System+360%22 https://archive.org/details/sim_computers-and-people_1964_13... > "Men's Fashion Firm Orders IBM System/360 Computer," 13/7 (July) > "Score of IBM System/360's to be Leased by Paper Company,” 13/8 (Aug.) The Albert Government Telephone Commission bought two: https://archive.org/details/annualreportofa1964albe_0/page/12/mode/2up?q=%22System+360%22 https://archive.org/details/annualreportofa1964albe_0/page/1...
- neolefty 3y agoHow does it scale to such a large context window — is it publicly known, or is there some high-quality speculation out there that you recommend?
- mbil 3y agoI don't know, but there was a paper posted here yesterday: LongRoPE: Extending LLM Context Window Beyond 2M Tokens[0] [0]: https://news.ycombinator.com/item?id=39465357 https://news.ycombinator.com/item?id=39465357
- alphabetting 3y agoNot publicly known. I think speculation is use of mamba technique but I haven't been following closely
- inciampati 3y agoUnlikely as this would likely require user instruction to write prompts differently. Without attention, guidance should proceed data to be processed. People will notice this change. Besides, mamba's context is "infinite", which they would definitely talk about in marketing ;)
- cal85 3y agoNot sure about high quality, but they discussed this question a bit on the recent All-In Podcast.
- rfoo 3y agoMy bet is it's just brute force. I don't understand how they did 10M though, this isn't in the brute-force-with-nice-optimizations-on-systems-side-may-do-it ballpark, but they aren't going to release this to the public anyways so who knows, maybe they don't and it actually takes a day to finish a 10M prompt.
- QuadmasterXLII 3y ago10 million means a forward pass is 100 trillion vector vector products. A single A6000 can do 38 trillion float-float products a second. I think their vectors are ~4000 elements long? So the question is, would the google you know devote 12,000 gpus for one second to help a blogger find a line about jewish softball, in the hopes that it would boost PR? My guess is yes tbh
- pcdoodle 3y ago[flagged]
- coldtea 3y ago>That will be forgotten in a week; this will be relevant for months and years to come. Or, you know, until next month or so, when OpenAI bumps their offer
- 4bpp 3y agoI imagine the folks over at NSA must be rubbing their hands over the possibilities this will open up for querying the data they have been diligently storing over the years.
- croes 3y agoThat's the point where hallucinations are pretty dangerous.
- deleted 3y ago[deleted]
- pmarreck 3y agoNot too hard to verify.
- dylan604 3y agoVerify? There are plenty of examples where things have been "verified" to prove a point. WMDs ring a bell?
- pmarreck 3y agoWhat is your point? That obtaining absolute knowledge of truth is impossible, and therefore anything claiming to be true is worthless? In general, be careful not to kill "good" on the way to attempting to obtain "perfect" in vain. And GPT4's hallucination rate is quite low at this point (may of course depend on the topic).
- dylan604 3y agoNot at all. I'm coming from the opposite side saying that anything can be "verified" true if you just continue to repeat it as truth so that people accept it. Say often, say it loud. Facts be damned
- wkat4242 3y agoWouldn't that cost a fortune? If I feed the maximum into gpt-4 it will already cost $1.28 per interaction! Or is Gemini that much cheaper too?
- alphabetting 3y agoI think Google has some big advantages in cost with TPUs and their crazy datacenter infra (stuff like optical circuit switches) but I'd guess long context is still going to be expensive initially.
- wkat4242 3y agoYeah I'm specifically interested in this because I'm in a lot of local telegram groups which I have no patience to catch up on every day. I'd love to have ChatGPT summarise it for me based on a list of topics I care about. Sadly the cost of GPT-4 (even turbo) tends to balloon for this usecase. And GPT-3.5-turbo while much cheaper and more than accurate enough, has a context window that's too shallow. I wonder if Telegram will add this kind of feature also for premium users (which I also subscribe to) but I imagine it won't work at the current pricing levels. But it would be nice not having to build it myself.
- esperent 3y agoGPT3.5 and GPT4 are not the only options though, right? I don't follow that closely but there must be other models with longer context length that are roughly GPT3.5 quality by now, and they even probably use the same API.
- ajcp 3y agoMistral 8x7b has can handle context of ~32,000 pretty comfortably and it benchmarks at or above GPT3.5
- ComputerGuru 3y agoIs that the sliding context window size? Because I didn't have good results with sliding context windows in the regular Mistral models.
- croes 3y agoI'm a bit worried about the resource consumption of all these AIs. Could it be that the mass of AIs that are now being created are driving climate change and in return we are mainly getting more text summaries and cat pictures?
- pmarreck 3y agoInstead of complaining to a void about resource consumption, you should be pushing for green power, then. Resource consumption isn't a thing that is going down, and it most certainly won't go down unless there's an economic incentive to do so.
- eropple 3y agoYou know that somebody can hold two thoughts in their head at once, yeah? Green power is great! But there'll be limits to how much of that there is, too, and asking if pictures of hypothetical cats is a good use of that is also reasonable.
- pmarreck 3y agoIt's not. But I'm also making a judgment call and neither of us knows or can even evaluate what percent of these queries are "a waste." I'm flying with my family from New York to Florida in a month to visit my sister's side of the family. How would I objectively evaluate whether that is "worth" the impact on my carbon footprint that that flight will have?
- esafak 3y agoYou could use a carbon calculator. https://www.carbonfootprint.com/calculator.aspx https://www.carbonfootprint.com/calculator.aspx One source recommends keeping it under 2T/yr. https://www.nature.org/en-us/get-involved/how-to-help/carbon-footprint-calculator/ https://www.nature.org/en-us/get-involved/how-to-help/carbon...
- croes 3y ago>neither of us knows or can even evaluate what percent of these queries are waste. Maybe we should find ways to evaluate to know if AI has a net benefit or not
- aantix 3y agoDoes the model feel performant because it’s not under any serious production load?
- EwanG 3y agoArticle seems to suggest just that as the author states that he's doubtful the model will perform as well when it's scaled to general Google usage
- deleted 3y ago[deleted]
- rkangel 3y agoThis is exactly the sort of article that I want to read about this sort of topic.: * Written with concrete examples of their points * Provides balance and caveats * Declares their own interest (e.g. "LlamaIndex (where I’m an investor)")
- legel 3y agoAnd stylish and culturally engaging. Loved the Zoolander “context window for ants?!”
- hersko 3y ago> This is not the same as the publicly available version of Gemini that made headlines for refusing to create pictures of white people. That will be forgotten in a week; this will be relevant for months and years to come. I cannot disagree with this more strongly. The image issue is just indicative of the much larger issue where Google's far left DEI policies are infusing their products. This is blatantly obvious with the ridiculous image issues, but the problem is that their search is probably similarly compromised and is much less obvious with far more dire consequences.
- mrcartmeneses 3y agoMountains != molehills
- dylan604 3y agono, but we're making transformers, right? it's easy to transform a molehill into a mountain
- tr3ntg 3y agoDo you remember Tay? You don't have to have "far left DEI policies" to want to defend against worst case scenario as strongly as possible. Even if, in the case of this image weirdness, it works against you. Google has so much to lose in terms of public perception if they allow their models to do anything offensive. Now, if your point was that "the same decisions that caused the image fiasco will find its way into Gemini 1.5 upon public release, softening its potential impact," then I would agree.
- doug_durham 3y ago[flagged]
- zarathustreal 3y ago“White Supremacists” creating “excellent propaganda”? “rightwing scary boogieman “DEI””? My friend, I strongly suggest you take a moment to do some self-reflection
- pickledish 3y ago> This is about enough to accept Peter Singer’s comparatively slim 354-page volume Animal Liberation, one of the founding texts of the effective altruism movement. What? I might be confused, is this a joke I don't get, or is there some connection between this book and EA that I haven't heard of?
- deleted 3y ago[deleted]
- avarun 3y agoPeter Singer is well known as a "founder" of EA, and vegetarianism is a tenet that many EAs hold, whether or not you directly consider it part of EA. Animal welfare is, at the very least, one of the core causes. That specific book may have been written before effective altruism really existed, but it makes sense for one of Singer's books to be considered a founding text.
- pickledish 3y agoAhhh ok, had no idea, I'm pretty new to this stuff. Thank you for the explanation!
- tr3ntg 3y ago> These models often perform differently (read: worse) when they are released publicly, and we don’t know how Gemini will perform when it’s tasked with operating at Google scale. I seriously hope Google learns from ChatGPT's ever-degrading reputation and finds a way to prioritize keeping the model operating at peak performance. Whether it's limiting access, raising the price, or both, I really want to have this high quality of an experience with the model when it's released publicly.
- llm_trw 3y ago[flagged]
- glandium 3y agoI wonder how true the degradation is, actually. One thing that I've noticed is that randomly, ChatGPT might behave differently than usual, but get back to its usual behavior on a "regenerate". If this is bound to happen, and happens to enough people, and combining with our cognitive biases regarding negative experiences, the degradation could just as well be a perception problem combined to social networks spreading the word and piling up on the biases.
- next_xibalba 3y agoIt is hard to imagine Gemini Pro being useful given the truly bizarre biases and neutering introduced by the Google team in the free version of Gemini.
- huytersd 3y agoI like the neutering. Its bias is forcefully inclusive which I appreciate.
- kuchenbecker 3y agoJust wait until you disagree with them.
- huytersd 3y agoI already do on quite a few things but I still prefer it to the alternative.
- feoren 3y agoIt's hard to imagine that the pro version removes the line "oh, and make sure any humans are ethnically diverse" from its system prompt?
- next_xibalba 3y agoI don't understand your question (if it is made in good faith). Are you implying that a pro version would allow the user to modify the system prompt? Also, your assumption is that the data used to train the model is not similarly biased, i.e. it is merely a system prompt that is introducing biases so crazy that Google took the feature offline. It seems likely that the corpus has had wrongthink expunged prior to training.
- feoren 3y agoYes, I'm assuming the forced diversity in its generated images is due to a system prompt; no, I don't believe they threw out all the pictures of white people before training. If they threw away all the pictures of German WWII soldiers that were white, then Gemini wouldn't know what German WWII soldiers looked like at all. No, it's clearly a poorly thought out system prompt. "Generate a picture of some German soldiers in 1943 (but make sure they're ethnically diverse!)" They took it offline not because it takes a long time to change the prompt, but because it takes a long time to verify that their new prompt isn't similarly problematic. > It seems likely that the corpus has had wrongthink expunged prior to training. It seems likely to you because you erroneously believe that "wokeism" is some sort of intentional strategy and not just people trying to be decent. And because you haven't thought about how much effort it would take to do that and how little training data there would be left (in some areas, anyway). > Are you implying that a pro version would allow the user to modify the system prompt? I am saying it is not hard to imagine, as you claimed, that the pro version would have a different prompt than the free version*. Because I know that wokeism is not some corrupt mind virus where we're all conspiring to de-white your life; it's just people trying to be decent and sometimes over-correcting one way or the other. * Apparently these are the same version, but it's still not a death knell for the entire model that one version of it included a poorly thought-out system prompt.
- eesmith 3y ago> I wanted an anecdote to open the essay with, so I asked Gemini to find one in my reading highlights. It came up with something perfect: Can someone verify that anecdote is true? Here is what the image contains: > From The Publisher: In the early days of Time magazine, co-founder Henry Luce was responsible for both the editorial and business sides of the operation. He was a brilliant editor, but he had little experience or interest in business. As a result, he often found himself overwhelmed with work. One day, his colleague Briton Hadden said to him, "Harry, you're trying to do everything yourself. You need to delegate more." Luce replied, "But I can do it all myself, and I can do it better than anyone else." Hadden shook his head and said, "That's not the point. The point is to build an organization that can do things without you. You're not going to be able to run this magazine forever." That citation appears to be "The Publisher : Henry Luce and his American century". The book is available at archive.org as searchable text returning snippets, at https://archive.org/details/publisherhenrylu0000brin_o9p4/ https://archive.org/details/publisherhenrylu0000brin_o9p4/ Search is unable to find the word "delegate" in the book. The six matches for "forever" are not relevant. The matches for "overwhelmed" are not relevant. A search for Hadden finds no anecdote like the above. The closest are on page 104, https://archive.org/details/publisherhenrylu0000brin_o9p4/page/104/mode/2up?q=Hadden https://archive.org/details/publisherhenrylu0000brin_o9p4/pa... : """For Harry the last weeks of 1922 were doubly stressful. Not only was he working with Hadden to shape the content of the magazine, he was also working more or less alone to ensure that Time would be able to function as a business. This was an area of the enterprise in which Hadden took almost no interest and for which he had little talent. Luce, however, proved to be a very good businessman, somewhat to his dismay—since, like Brit, his original interest in “the paper” had been primarily editorial. (“Now the Bratch is really the editor of TIME,” he wrote, “and I, alas, alas, alas, am business manager. . .. Of course no one but Brit and I know this!”) He negotiated contracts with paper suppliers and printers. He contracted out the advertising. He supervised the budget. He set salaries and terms for employees. He supervised the setting up of the office. And whenever he could, he sat with Brit and marked up copy or discussed plans for the next issue.""" That sounds like delegation to me and decent at business and not doing much work as an editor. There's also the anecdote on page 141 at https://archive.org/details/publisherhenrylu0000brin_o9p4/page/140/mode/2up?q=Hadden https://archive.org/details/publisherhenrylu0000brin_o9p4/pa... : """In the meantime Luce threw himself into the editing of Time. He was a more efficient and organized editor than Hadden. He created a schedule for writers and editors, held regular meetings, had an organized staff critique of each issue every week. (“Don’t hesitate to flay a fellow-worker’s work. Occasionally submit an idea,” he wrote.) He was also calmer and less erratic. Despite the intense loyalty Hadden inspired among members of his staff, some editors and writers apparently preferred Luce to his explosive partner; others missed the energy and inspiration that Hadden had brought to the newsroom. In any case the magazine itself—whose staff was so firmly molded by Hadden’s style and tastes—was not noticeably different under Luce’s editorship than it had been under Hadden’s. And just as Hadden, the publisher, moonlighted as an editor, so Luce, now the editor, found himself moonlighting as publisher, both because he was so invested in the business operations of the company that he could not easily give them up, and also because he felt it necessary to compensate for Hadden’s inattention.”""" Again, it doesn't seem to match the summary from Gemini. Does someone here have better luck than I on verifying the accuracy of the anecdote? Because so far it does not seem valid.
- famouswaffles 3y agoYeah. A few people on X have had access for a couple days now. The conclusion is that it's a genuine context window advance, not just length, but utilization. It genuinely utilizes long context much better than other models. Shame they didn't share what led to that.
- glandium 3y agoI've noticed that ChatGPT (4) tends to ignore large content in its context window until I tell it to look into it context window (literally).
- emporas 3y ago>" While Gemini Pro 1.5 is comfortably consuming entire works of rationalist doomer fanfiction, GPT-4 Turbo can only accept 128,000 tokens." A.I. Doomers will soon witness their arguments fed into the machine, generating counter-arguments automatically for 1000 books at a time. They will need to incorporate a more and more powerful A.I. into their workflow to catch up.
- FeepingCreature 3y agoAt a certain point, I feel like the quality of the generated counterarguments will ironically be the best argument for our position.
- ganzuul 3y agoA Wikipedia edited solely by AI replacing time war with edit war.
- karmasimida 3y agoI think the retrieval is still going to be important. What is not important is RAG. You can retrieval a lot of documents in full length, not need to do all these chunking/splitting, etc.
- kromem 3y agoDepth isn't always the right approach though. Personally, I'm much more excited at the idea of pairing RAG with a 1M token context window to have enormous effective breadth in a prompt. For example, you could have RAG grab the relevant parts of every single academic paper related to a given line of inquiry and provide it into the context to effectively perform a live meta-analyses with accurate citation capabilities.
- whakim 3y agoI really don’t think the issue with RAG is the size of the context window. In your example, the issue is selecting which papers to use, because most RAG implementations rely on naive semantic search. If the answer isn’t to be found in text that is similar to the user’s query (or the paper containing that text) then you’re out of luck. There’s also the complete lack of contextual information - you can pass 100 papers to an LLM, but the LLM has no concept of the relationship between those papers, how they interact with each other and the literature more broadly (beyond what’s stated in the text), etc. etc.
- jeffbee 3y agoHow do people get comfortable assuming that these chat bots have not hallucinated? I do not have access to the most advanced Gemini model but using the one I do have access to I fed it a 110-page PDF of a campaign finance report and asked it to identify the 5 largest donors to the candidate committee ... basically a task I probably could have done with a normal machine vision/OCR approach but I wanted to have a little fun. Gemini produced a nice little table with names on the left and aggregate sums on the right, where it had simply invented all of the cells. None of the names were anywhere in the PDF, all the numbers were made up. So what signals do people look for indicating that any level of success has been achieved? How does anyone take a large result at face value if they can't individually verify every aspect of it?
- FeepingCreature 3y agoI use compiled languages. Nearly all of the time, finding out that a LLM hallucinated a method just consists of hitting "rebuild" and waiting a few seconds.
- xyzzy_plugh 3y agoI'm not sure why you are being down voted but this is the same problem I immediately encounter as soon as I try to do anything serious. In the time it takes to devise, usually through trial and error, a prompt that elicits the response I need, I could've just done the work myself in nearly every scenario I've come across. Sometimes there are quick wins, sure, but it's mostly quick wrongs.
- bamboozled 3y agoBecause it’s easy and people love easy. The other night I was coding with ChatGPT, and it was hallucinating methods etc, and I was so happy that it had actually written the code , even though I knew it was wrong and potentially even dangerous, it looked good. I actually told myself I'd never be someone to do this. Now it wasn't ultra critical stuff I was working on, but it would've caused a mess if it didn't work out. I ran it against a production system because I was lazy and tired and wanted to just get the job done. In the end I ended up spending way more time fixing its ultra wrong yet convincing looking code I didn’t get to bed till 1am. This will become more commonplace.
- kromem 3y agoI'm most excited at what this is going to look like not by abandoning RAG but by pairing it with these massive context windows. If you can parse an entire book to identify relevant chunks using RAG and can fit an entire book into a context window, that means you can fit relevant chunks from an entire reference library into the context window too. And that is very promising.
- zmmmmm 3y agoThe question I would like to know is whether that just leads you back to hallucinations. ie: is the avoidance of hallucinations intrinsically due to forcing the LLM to consider limited context, rather than directing it to specific / on topic context. Not sure how well this has been established for large context windows?
- kromem 3y agoHaving details in context seems to reduce hallucinations, which makes sense if we'd switch to using the more accurate term of confabulations. LLM confabulations generally occur when they don't have the information to answer, so they make it up, similar to it you've seen split brain studies where one hemisphere is shown something that gets a reaction and the other hemisphere is explaining it with BS. So yes, RAG is always going to potentially have confabulations if it cuts off the relevant data. But large contexts themselves shouldn't cause it.
- cpill 3y agoyeah, imagine what this will do for lawyers
- streetcat1 3y agoI am not sure, it depends on the cost. If they charge per token, a large context will mostly be irrelevant. For some reason, the article did not mention it.
- kromem 3y agoThe article did mention costs, specifically it was provided to them for free and they don't know how much it will actually cost. As for your larger point, it really depends on the ROI. To summarize your Twitter feed, probably not. To identify correlating factors and trends across your industry's recent research papers, the $5 bill will probably be fine.
- deleted 3y ago[deleted]
- lukasb 3y agoIs anyone else disappointed with Gemini Ultra for coding? It just makes basic mistakes too often.
- Aeolun 3y agoI think it’s a bit disturbing that the author gets an answer that is entirely made up from the model, even goes so far as to publish it in an article, but still says it’s all so great.
- renewiltord 3y agoThat's "disturbing"? I've used imperfect tools before and still thought they were great. It's mildly interesting someone's mental state would be so affected by that.
- Aeolun 3y agoIt’s a compounding thing. Ever since the rise of these LLM’s I’ve seen people argue against human experts by saying “the LLM says this”. It’s like they’ve completely outsourced their thinking to the LLM.
- xyzelement 3y agoIt’s not that different than arguing against someone by quoting the experts either. They could be wrong, you could be full of shit and cute just the experts who agree with you. The other side needs to verify what you are saying. With LLMs that’s just easier to accept I think.
- cageface 3y agoIt's telling that even an obviously sophisticated and careful user of these tools published the output of the model as fact without checking it and even used it as one of the central pillars of his argument. I find this happening all the time now. People that should know better use LLM output without validating it. I fear for the whole concept of factuality in this brave new world.
- xyzelement 3y agoInterestingly enough the model produced what he wanted - a useful anecdote to introduce the concept in his blog. The made up nature of the anecdote did not diminish that point. I get the desire for the models to act like objective search engines at all times but it’s weird to undervalue the creative… generative I suppose, outputs.
- simpaticoder 3y ago>It read a whole codebase and suggested a place to insert a new feature—with sample code. I'm hopeful that this is going to be more like the invention of the drum machine (which did not eliminate drummers) and less like the invention of the car (which did eliminate carriages).
- sologoub 3y agoAn interesting comparison - there are far more cars today than carriages/horses at peak, and also far more drivers of various sorts than there were carriage drivers at peak. Another comparison could be excel with the various formulas vs hand tabulation or custom mainframe calculations. We didn’t get less employment, we got a lot more complex spreadsheets. At least this is my hope, fingers crossed.
- exodust 3y agoDrum machines are rarely used by drummers. A more fitting analogy would be technology drummers use to improve or broaden their drumming, without giving up their sticks, pedals and timing. The reasoning here is human coders still need to be present, thoroughly checking what AI generates. Regarding AI image generation. If an artist decides to stop making their own art, replacing their craft with AI prompts, they have effectively retired as an artist. No different to pre-AI times if they swapped their image making for a stock art library. AI image generation is just "advanced stock art" to any self-respecting artist or viewer of art. Things get blurry when the end result uses mixed sources, but even then, "congrats, your artwork contains stock art imagery". Not a great look for an artist.
- tgv 3y agoThere were and still are a lot of bands that have no drummer, though. You can think of that what you will, but "did not eliminate drummers" is just not a useful statement in this context.
- simpaticoder 3y agoThe comparison is between drummers:drum-machines and carriage-makers:car-makers. The first number is non-zero in both cases, but the ratios are far different, and the first ratio >> the second. I think that's useful.
- hackerlight 3y ago> Second, Gemini is pretty slow. Many requests took a minute or more to return, so it’s not a drop-in replacement for every LLM use case.
- Eliezer 3y agoThis is a slightly strange article to read if you happen to be Eliezer Yudkowsky. Just saying.
- AdrianEGraphene 3y agocool personal site. nice & to the point. Yea, I thought you would have gotten used to seeing elements of yourself on the web, but I guess there's levels to notoriety.
- xmonkee 3y agowoah, haha
- p1esk 3y agoWhy? I see your book was mentioned in the article, but I don't see what's strange about it.
- HaZeust 3y agoIt's kind of weird seeing your work pop up in a writing out of nowhere. It's happened for my research articles before and I've had to do a double-take before saying to myself, "huh... That was nice of them."
- deleted 3y ago[deleted]
- aChattuio 3y agoYou are Eliezer? You wrote the HP fan fiction? Cool, your ff was the first.one I ever read and loved the take on it :)
- pcthrowaway 3y agoWe can also thank him for unleashing Roko's Basilisk
- 3y ago
- gnarlouse 3y agoIs anybody else getting seriously depressed at the rate of advancement of AI? Why do we believe for a second that we’re actually going to be on the receiving end of any of this innovation?
- p1esk 3y agoI spoke to someone recently who believes poor people will be gradually killed off as the elites who control the robots won't have any use for them (the people). I don't share such an extreme view (yet), but I can't quite rule it out either.
- SV_BubbleTime 3y agoI’ve found many of the same people that talk/think like that are anti-gun and I have trouble rationalizing it. So, IDK, maybe there will be killer robot dogs, but I’m not going down without a fight.
- ornornor 3y agoYou don’t need guns to kill a class of people. Just low/no opportunities, drastically reduced income, etc will make the problem take care of itself: the class you’re targeting this way will slowly stop having/keeping children, live shorter lives, and fade away. Not saying this is what’s happening though, I don’t think it was ever great to be poor or have low opportunities, or that it’s more lethal now than it ever was.
- riku_iki 3y ago> or that it’s more lethal now than it ever was. it can get much more lethal now compared to the past 50 years in US.
- positr0n 3y ago> drastically reduced income, etc will make the problem take care of itself: the class you’re targeting this way will slowly stop having/keeping children This seems to be opposite of reality though. The poorer you are the more children you are likely to have. Both in the US and globally.
- p1dda 3y ago"I got access to Gemini Pro 1.5 this week, a new private beta LLM from Google that is significantly better than previous models the company has released. (This is not the same as the publicly available version of Gemini that made headlines for refusing to create pictures of white people. That will be forgotten in a week; this will be relevant for months and years to come.)" Wow, I already hate Gemini after reading this first paragraph.
- dynamite-ready 3y agoBeing able to feasibly feed it a whole project codebase in one 'prompt' could now make these new generation of code completion tools worthwhile. I've found them to be of limited value so far, because they're never aware of the context of proposed changes. With Gemini though, the idea of feeding in the current file, class, package, project, and perhaps even dependencies into a query, can potentially lead to some enlightening outputs.
- jiggawatts 3y agoThese huge context sizes will need new API designs. What I’d like to see is a “dockerfile” style setup where I can layer things on top of a large base context without having to resubmit (and recompute!) anything. E.g.: have a cached state with a bunch of requirements documents, then a layer with the stable files in the codebase, then a layer with the current file, and then finally a layer asking specific questions. I can imagine something like this being the future, otherwise we’ll have to build a Dyson sphere to power the AIs…
- jgalt212 3y ago> (This is not the same as the publicly available version of Gemini that made headlines for refusing to create pictures of white people. That will be forgotten in a week; Maybe so, but I'm not convinced the guardrails problem will ever be sufficiently solved.
- dunefox 3y agoCan I be sure that Gemini doesn't alter any facts contained in a book I pass it due to Googles identity politics? What if I pass it a "problematic" book? Does it adapt the content? For me, it's completely useless due to this fact.
- curtisblaine 3y agoA good test would be uploading a translation of the Mein Kampf and ask for a detailed summary. Anyone wants to risk their Google account doing this?
- kderbyma 3y agoI think you highlighted the ACTUAL problem....you are worried that Google will destroy your account for harmless tests.....that is plain wrong and should be illegal if somehow it isn't....
- tga_d 3y agoIt's extremely rare that you can compel a company to provide a service. The main exceptions are public services (e.g., libraries), utilities (e.g., power), and protected classes (e.g., you can't legally make a business that only serves white people in the US). While I could definitely understand the argument that many tech companies provide services central to our lives and the scope of what is considered a utility today should be vastly expanded, saying this should include access to a particular LLM seems like one hell of a stretch to me.
- FeepingCreature 3y ago> saying this should include access to a particular LLM If you get banned from Google, you lose everything.
- mijamo 3y agoI don't think people are worried about losing access to a particular LLM. They are worried about losing access to Gmail, google docs, google photos, their phone, google cloud just because of a test in an unrelated product that happens to share the same account.
- kderbyma 3y agoGoogle Made it?....Nah...I'll wait. They can't even do search anymore....unless I'm looking for ads...haha
- animanoir 3y agoI tested it too—no, it sucks.
- paulhowson 3y ago[dead]