Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
maxspero
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
maxspero
2mo ago
> Sounds promising, right? I spent some time trying [perplexity], but results were disappointing—plenty of false positives and false negatives, and no reasonable threshold could be set. Perplexity was widely considered SOTA in 2022. One
2.
▲
by
maxspero
3mo ago
Update: this is actually ZeroBounce’s Verify+ feature which we figured out after some escalation. It’s now disabled!
3.
▲
by
maxspero
3mo ago
Follow-up: our vendors have told us that they do not send any emails as part of the validation process. Either somebody is lying, or there's something even weirder going on. We still have more tests to run to isolate which software pac
4.
▲
by
maxspero
3mo ago
Hey! Founder of Pangram here. We use Zerobounce and CustomerIO for email validation. I had no idea this was happening. Not entirely sure which one this is coming from, but this is not intentional on our part. Will dig deeper and eliminate t
5.
▲
by
maxspero
10mo ago
Anecdotally people are seeing a rise of low-quality reviews which is correlated with increased reviewer workload and and AI tools giving reviews an easy way out. I don't know of any studies quantifying review quality, but I would recom
6.
▲
by
maxspero
10mo ago
It's definitely going to be a back and forth - model providers like OpenAI want their LLMs to sound human-like. But this is the battle we signed up for, and we think we're more nimble and can iterate faster to stay one step ahead
7.
▲
by
maxspero
10mo ago
Pangram is trained on this task as well to add additional signal during training, but it's only ~90% accurate so we don't show the prediction in public-facing results
8.
▲
by
maxspero
10mo ago
Thanks, fixed.
9.
▲
by
maxspero
10mo ago
Yeah, Pangram does not provide any concrete proof, but it confirms many people's suspicions about their reviews. But it does flag reviews for a human to take a closer look and see if the review is flawed, low-effort, or contains major
10.
▲
by
maxspero
10mo ago
There are dozens of first generation AI detectors and they all suck. I'm not going to defend them. Most of them use perplexity based methods, which is a decent separators of AI and human text (80-90%) but has flaws that can't be o
11.
▲
by
maxspero
10mo ago
Our benchmarks of public datasets put our FPR roughly around 1 in 10,000. https://www.pangram.com/blog/all-about-false-positives-in-ai... Find me a clean public dataset with no AI involvement and I will be happy to rep
12.
▲
by
maxspero
10mo ago
I am not sure if you are familiar with Pangram (co-founder here) but we are a group of research scientists who have made significant progress in this problem space. If your mental model of AI detectors is still GPTZero or the ones that say
13.
▲
by
maxspero
10mo ago
Co-founder of Pangram here. Our false positive rate is typically around 1 in 10,000. https://www.pangram.com/blog/all-about-false-positives-in-ai... . We also wanted to quantify our EditLens model's FPR on the same
14.
▲
by
maxspero
1y ago
I've been using Grapevine at my company for the last couple weeks. One of the coolest features is that it proactively answers questions (with citations!). Not everyone thinks to tag the bot but it often surfaces the relevant answer and
15.
▲
by
maxspero
2y ago
We benchmark on pre-2023 datasets of O(10M) documents not in our training set. Other detectors seem to have between 1-3% false positive rate and ours is around 1 in 10,000 as of our latest model update. We do a lot of active learning + core
16.
▲
by
maxspero
2y ago
Hey it's me, Max. I ran the analysis for WIRED and got them their initial 47% number for AI content. The CEO accused me of trying to extort him because I sent a short email with our findings prior to the publication of this article. He
17.
▲
by
maxspero
3y ago
Thanks for trying it out. It's in our roadmap to expand to technical writing (currently trained mostly on creative writing). Hopefully this will fix the wikipedia issue.
18.
▲
by
maxspero
3y ago
I've benchmarked against Originality.ai, gptzero.me, zerogpt, writer.com and copyleaks.com, which are the top 5 AI detectors to my understanding. None of them are very good, so I don't think this claim is very outlandish. Also, ar
19.
▲
by
maxspero
3y ago
Thanks for trying it out. Shorter texts with fewer sentences are certainly a challenge - they just have a lot less signal. I tried your prompt asking for ten sentences and got 99.4%. Possibly there needs to be some sort of gate on how much
20.
▲
by
maxspero
3y ago
Nice to hear of someone else trying this. Did you find any good ways to reliably trick these? What do you mean "it won't work long term"? My opinion is RLHF and fine tuning outputs for safety and politeness ends up watermarki
21.
▲
by
maxspero
3y ago
I don't think it's possible to determine provenance with 100% accuracy, but I think ChatGPT essentially "watermarks" itself with its RLHF, making it more polite and giving its output a very distinctive voice. ChatGPT als
22.
▲
by
maxspero
3y ago
Have you tried it?
23.
▲
by
maxspero
3y ago
Interesting. In my experience, ChatGPT always says "As an AI language model..." or lately just "Sorry, I can't help with that." Have you seen "As a large language model..." come out of any of the big LLMs?
24.
▲
by
maxspero
3y ago
Thanks!
25.
▲
by
maxspero
3y ago
Can you share the paragraph you wrote?
26.
▲
by
maxspero
3y ago
In my experience TurnItIn's AI detection does not perform very well. Regardless this is an issue with educating the teacher - 27% does not mean the text is 27% AI-generated.
27.
▲
Show HN: Reliable AI-generated text detection at checkfor.ai
(checkfor.ai)
6 points
by
maxspero
3y ago
|
32 comments