Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Gcam
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Show HN: Artificial Analysis tool to create custom benchmarks for any use case
(artificialanalysis.ai)
11 points
by
Gcam
1mo ago
|
0 comments
2.
▲
by
Gcam
1mo ago
Hey! George from the Artificial Analysis team here. We published an update today that does result in a change of the order, Qwen3.8 Max to second rather than first. The methodology change was an already planned upgrade to our equality check
3.
▲
by
Gcam
9mo ago
Artificial Analysis | https://artificialanalysis.ai | Full Stack Engineer, ML Engineer, Member of Technical Staff | Onsite in San Francisco or in Australia/New Zealand | Competitive Salary + Equity Artificial Analysis is an
4.
▲
Show HN: Stirrup – A lightweight and customizable foundation for building agents
(github.com)
2 points
by
Gcam
9mo ago
|
0 comments
5.
▲
by
Gcam
1y ago
Artificial Analysis | https://artificialanalysis.ai | Full Stack Engineer, ML Engineer, Member of Technical Staff | Onsite in San Francisco or in Australia/New Zealand | Competitive Salary + Equity Artificial Analysis is an
6.
▲
by
Gcam
1y ago
Artificial Analysis | https://artificialanalysis.ai | Full Stack Engineer & ML Engineer | Onsite in San Francisco or in Australia/New Zealand | Competitive Salary + Equity Artificial Analysis is an independent AI benchm
7.
▲
MicroEvals – Easily run vibe checks against models
(artificialanalysis.ai)
3 points
by
Gcam
1y ago
|
0 comments
8.
▲
by
Gcam
2y ago
Artificial Analysis | https://artificialanalysis.ai | Full Stack Engineer & ML Engineer | Onsite in San Francisco or in Australia/New Zealand | Competitive Salary + Equity Artificial Analysis is an independent AI benchm
9.
▲
by
Gcam
2y ago
Data in this tool is from https://artificialanalysis.ai/ on October 13 2024 and so is a little of out date. This page has up to date information of all models and providers: https://artificialanalysis.ai/lea
10.
▲
by
Gcam
2y ago
Artificial Analysis | https://artificialanalysis.ai | Senior Full Stack Software Engineer | Onsite in San Francisco | Competitive salary + Equity Artificial Analysis is an independent benchmarking, evaluation and insights provid
11.
▲
by
Gcam
2y ago
Artificial Analysis | https://artificialanalysis.ai | Senior AI Analyst & Senior Software Developers | Remote or Hybrid (USA, San Francisco preferred, or Australia) | Competitive salary + Equity We're seeking a Senior A
12.
▲
by
Gcam
2y ago
See here for our TTFT metric benchmarks: https://artificialanalysis.ai/models/llama-3-1-instruct-70b/...
13.
▲
by
Gcam
2y ago
this page is probably our most comparable to thefastest (which is cool, more benchmarks is better): https://artificialanalysis.ai/leaderboards/providers We also have pricing, long/medium/short prompt lengths
14.
▲
by
Gcam
2y ago
What is your view on the E-2 visa for startup founders and likelihood of success (with an investment of ~50-100k for software business)?
15.
▲
by
Gcam
3y ago
As part of our benchmarking of Groq we have asked Groq regarding quantization and they have assured us they are running models at full FP-16. It's a good point and important to check. Link to benchmarking: https://artificial
16.
▲
by
Gcam
3y ago
Groq's API performance reaches close to this level of performance as well. We've benchmarked performance over time and >400 tokens/s has sustained - can see here https://artificialanalysis.ai/models/mi
17.
▲
From GPT-4 to Mistral 7B, there is a 300x range in the cost of LLM inference
(twitter.com)
2 points
by
Gcam
3y ago
|
0 comments
18.
▲
Show HN: LLM Benchmarks Leaderboard with 60 model and API host combinations
(artificialanalysis.ai)
3 points
by
Gcam
3y ago
|
1 comments
19.
▲
Mistral API reduces time to first token by 10x (only place for Mistral Medium)
(twitter.com)
4 points
by
Gcam
3y ago
|
0 comments
20.
▲
240 Tokens/s achieved by Groq's custom chips on Lama 2 Chat (70B)
(twitter.com)
5 points
by
Gcam
3y ago
|
0 comments
21.
▲
by
Gcam
3y ago
Hey com2kid - if you're still there, we did end up adding boxplots to show variance. Can be seen on the models page https://artificialanalysis.ai/models and on each models page where you view hosts by clicking one of t
22.
▲
New GPT-4 Turbo (0125 Preview) slightly faster per initial benchmarks
(twitter.com)
2 points
by
Gcam
3y ago
|
0 comments
23.
▲
by
Gcam
3y ago
Definitely agree with your point on Claude Instant though. Much less than half the price, much higher throughput/speed for a relatively small quality decrease (varied by how 'quality' is measured, use-case)
24.
▲
by
Gcam
3y ago
Hi, we have this if you take a look at the models page ( https://artificialanalysis.ai/models ) and scroll down to 'Latency', and also on the API host comparison pages for each model (e.g. https://artific
25.
▲
by
Gcam
3y ago
We have Claude Instant on the models page: https://artificialanalysis.ai/models Can add it via the select at the top right of each card where it says '9 Selected' (below the highlight charts)
26.
▲
by
Gcam
3y ago
Model quality index methodology is as per this comment (can add perplexity using the dropdown): https://news.ycombinator.com/item?id=39014985#39017632 It's a combination of different quality metrics which have Perplexi
27.
▲
by
Gcam
3y ago
Thanks for the feedback and glad it is useful! Yes, agree might better representative of future use. I think a view of variance would be a good idea, currently just shown in over-time views - maybe a histogram of response times or a box an
28.
▲
by
Gcam
3y ago
Hi HN, Thanks for checking this out! Goal with this project is to provide objective benchmarks and analysis of LLM AI models and API hosting providers to compare which to use in your next (or current) project. Benchmark comparisons include
29.
▲
by
Gcam
3y ago
We have this (and other more detailed metrics) on the models page https://artificialanalysis.ai/models if you scroll down and for individual hosts if you click into a model (nav or click one of the model bars/bubbles)
30.
▲
by
Gcam
3y ago
Thanks! For Claude instant, select the dropdown on the top right of the card where it says '8 Selected' and can add it to the graphs. Thanks for the suggestions for adding Phi 2, Model.com as a host, can look into these!
More ›