Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
g023
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
1.
▲
by
g023
6d ago
And you don't think it is weird that the chinese models are the only ones embracing free markets, and open concepts, while the so called 'free' market on this side of the world tries to close everyone off from the technology
2.
▲
by
g023
7d ago
its probably more expensive to run, and I'm thinking 4.1 flash is a smaller more efficient model. You can always host your own. This is what they need to do to stay competitive.
3.
▲
by
g023
17d ago
I miss the old days of ICQ and just dragging a file onto the person you are sending your file and bam, done like dinner.
4.
▲
by
g023
1mo ago
All this performance at such small model sizes, why are the API fees so high for the AI monopolists on this side of the world?
5.
▲
by
g023
1mo ago
g023 Code - https://github.com/g023/g023code a deepseek v4 flash harness built specifically for that model, that uses Ollama for its vision support, and uses the built in web search capability natively. Created only in
6.
▲
by
g023
1mo ago
It says 'yes' where the others say 'no'. Good enough for me.
7.
▲
by
g023
1mo ago
Pure-Python AI coding agent powered by DeepSeek V4 Flash - Subagent-First architecture · Context is Currency · Terminal-native - Can connect Ollama models for Vision support - Uses the native web search built into DS v4 API and is built aro
8.
▲
by
g023
3mo ago
Slap the gpus in a car and offset the cost of ownership by supplying the grid for GPU power on the go. Either get paid in rebates or tokens. Contribute to a distributed training/inferencing network.
9.
▲
by
g023
3mo ago
Anything to close Pandora's box. "They" liked the eras they could control the communications, and therefore the narrative. Boomers on their last legs, question is, will the future undo the unjustness that was forced upon them
10.
▲
by
g023
3mo ago
gee I wonder how their models learned Chinese?
11.
▲
by
g023
3mo ago
How come people these days treat jobs like its a social gathering?
12.
▲
by
g023
3mo ago
A self-contained CUDA inference engine for LiquidAI/LFM2.5-8B-A1B (hybrid conv + GQA-attention MoE, 8.5B params, 1B active) targeting a single RTX 3060 (12 GB) using flash-decoding. MIT license.
13.
▲
by
g023
3mo ago
I wonder if "battelites" might be profitable. Like an pay-per-usage energy grid in space with battery backup that can beam power around to other satellites that might not have easy access to power, or have their power grids tempor
14.
▲
by
g023
3mo ago
China has a couple going at the moment.
15.
▲
by
g023
4mo ago
When did "hate the customer" become a thing?
16.
▲
by
g023
4mo ago
I use DeepSeek v4 flash with CoPilot and it works pretty good.
17.
▲
by
g023
4mo ago
If anyone is looking to hook it up to copilot, I made a proxy script to handle the connection a bit back that might be handy: https://gist.github.com/g023/c2bb7b540ffe64cee76023f18f6f936...
18.
▲
by
g023
5mo ago
Terminal-based chat application powered by *locally installed* llama.cpp, featuring an auto-managed server backend, reasoning modes, and 7 built-in filesystem tools for interactive AI assistance with a focus on only allowing read only agent
19.
▲
by
g023
5mo ago
We need more personal level AI solutions instead of so much corporate centered solutions.
20.
▲
Local Model Router: Ollama/OpenAI-compat bridges for local LLMs via llama.cpp
1 points
by
g023
5mo ago
|
0 comments
21.
▲
by
g023
5mo ago
I've started creating https://github.com/g023/localmodelrouter/ which offers Ollama like functionality but as a single .py file with minimal dependencies and more focus on letting llama.cpp handle the dirty w
22.
▲
by
g023
5mo ago
HarnessHarvester generates executable Python harnesses from natural language task descriptions, executes them in a sandboxed environment, reviews them with multi-faceted LLM judges, and repairs failures using branching strategies. It includ
23.
▲
by
g023
5mo ago
A single file, python based, minimal/recognizable dependencies, turboquant playground, barebones af, with some easy to access globals to experiment with at top of 'run_tquant.py'. Test model is a 1.77B model that I altered by
24.
▲
by
g023
5mo ago
I had some issues in the original, but had to jump away for a bit here to do some backups (weak). Anyways, I updated to make the necessary fixes, and also made some more tweaking values at top to play with and dialed in the params for the m
25.
▲
Show HN: Standalone TurboQuant KV Cache Inference
(github.com)
3 points
by
g023
6mo ago
|
4 comments
26.
▲
Show HN: An offline first focused agentic CLI application powered by Ollama
4 points
by
g023
6mo ago
|
0 comments
27.
▲
Show HN: G023's Agentic Chat with Memory and Python Power
(github.com)
1 points
by
g023
6mo ago
|
1 comments
28.
▲
by
g023
6mo ago
A sophisticated multi-level reasoning engine with agentic memory, tool integration, and user control modes. Built as a (primarily) single-file Python program using vanilla Python and local LLM integration. Powered by Ollama API and utilizes
29.
▲
by
g023
8mo ago
I made a smaller sized version https://huggingface.co/g023/Qwen3-8B-DMS-8x-4bit-NF4
30.
▲
by
g023
9mo ago
I'm thinking that the bubble will be the vortex caused by an abundance of power that becomes freely available locally due to the AI datacenters moving to space.
More ›