Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cgorlla
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
cgorlla
22d ago
Very likely, NYT confirmed it's releasing Friday.
2.
▲
by
cgorlla
22d ago
It's decent. How good depends on compute needs.
3.
▲
by
cgorlla
22d ago
It's fun! Also it's an interesting commentary on where the AI industry is as a whole.
4.
▲
Behaviorally fingerprinting Ox Alpha's provenance
(ctgt.ai)
41 points
by
cgorlla
22d ago
|
19 comments
5.
▲
by
cgorlla
1mo ago
Based on the reception of our last post we took folks' suggestions to run the new official build of V4 Flash and compare it to the preview that was released 5 days apart. The new build is added to https://playground.ctgt.ai&
6.
▲
by
cgorlla
1mo ago
Every post is actually scanned for LLM content as well.
7.
▲
by
cgorlla
1mo ago
A typographic Voight-Kampff test is pretty awesome.
8.
▲
DeepSeek V4-Flash-0731 is 12 pts more censored than preview (selectively)
(ctgt.ai)
3 points
by
cgorlla
1mo ago
|
0 comments
9.
▲
by
cgorlla
2mo ago
We discuss this in the writeup. While we expected this result, it is important for there to be data backing the claims, and an experimental setup that mirrors productions tasks is a useful tool for the conversations going on about this.
10.
▲
by
cgorlla
2mo ago
We're actually exploring the changes in the model geometry that cause it to comply or not comply with a given policy next, I think visual representations of that behavior would be interesting and perhaps elucidating. What you mention i
11.
▲
by
cgorlla
2mo ago
I guess the AI that wrote your comment for you also conflated the SFT step of the target domain with the political prompts, which, in the sentence you quoted, contradicts your original comment...
12.
▲
by
cgorlla
2mo ago
>We plan to test what happens with a Chinese teacher into a Chinese-lineage base like Qwen next. :)
13.
▲
by
cgorlla
2mo ago
The fact that you literally thought the examples were used in SFT in your last comment ago calls into question the utility of this conversation, notwithstanding the implication that those examples were used to improve…financial performanc
14.
▲
by
cgorlla
2mo ago
The examples you're talking about are not involved in the training process, so their number is irrelevant. As stated in the post, the goal of this work is to determine whether a teacher's unrelated behaviors are inherited by the s
15.
▲
by
cgorlla
2mo ago
You can try it yourself! https://playground.ctgt.ai
16.
▲
by
cgorlla
2mo ago
This is fixed
17.
▲
by
cgorlla
2mo ago
This is fixed.
18.
▲
by
cgorlla
2mo ago
Agreed, we find this to be an interesting reflection of societal values and norms inasmuch LLMs are.
19.
▲
by
cgorlla
2mo ago
Abliterated models certainly have their uses but they're not the default choice for most users or enterprises, and thus not the versions of those models most would interact with.
20.
▲
by
cgorlla
2mo ago
Agreed. It's fixed
21.
▲
by
cgorlla
2mo ago
You can see exactly what prompts we used and the results here: https://github.com/CTGT-Inc/lineage-eval/tree/main/data We found V4 Flash was significantly more censored than the baseline.
22.
▲
by
cgorlla
2mo ago
It's most likely to occur when distilling a Chinese model from a Chinese base. We plan to do compliance geometry analysis in the future to see what is structurally changing in the model when distillation causes it to start refusing or
23.
▲
by
cgorlla
2mo ago
Consider that LLMs are trained on the corpus of the internet, and (simplifying) consequently give the average answer of the internet. If the desired answer of the censorer is contradictory to this, then it requires additional training data
24.
▲
Show HN: Distilling DeepSeek into GPT-OSS doesn't transfer censorship. Try it
(ctgt.ai)
170 points
by
cgorlla
2mo ago
|
73 comments
25.
▲
by
cgorlla
9mo ago
>You're going to need an incredibly compelling sales pitch for me to send my data to an unknown vendor I agree! Our customers require on-prem deployments, though, so nothing is being sent to us outside their environment.
26.
▲
by
cgorlla
9mo ago
Glad you played around with it and that our tech worked.
27.
▲
by
cgorlla
9mo ago
SOTA results are a happy byproduct of the core mission of our approach, which is to enable the effective and simple translation of policy documents into a model without having to fine-tune and prompt engineer. This performance is somewhat u
28.
▲
by
cgorlla
9mo ago
We'll be back when the Holy War begins.
29.
▲
by
cgorlla
9mo ago
The product integrates as a layer on top of their existing models, serving as a policy-as-code layer so they don't have to fine-tune, prompt engineer etc. to get them up to par in their deployments as is standard now. One example that
30.
▲
by
cgorlla
9mo ago
I checked with the team and it may have been some temporary rate-limiting issue. We've rectified the results, it seems to be an isolated case. https://www.ctgt.ai/benchmarks
More ›