Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ipotapov
searching Neon…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
27 ms
·
1.
▲
US and Iran Exchange Attacks for First Time in About a Month
(bloomberg.com)
3 points
by
ipotapov
17d ago
|
0 comments
2.
▲
by
ipotapov
1mo ago
Your feature of speaking text from any app with a hotkey is a neat implementation for macOS. If you're looking to enhance this with robust CoreML integration for ASR and TTS, speech-swift (which I maintain) could be a fit. It offers a
3.
▲
A full offline voice agent in 1.2 GB of RAM on Android device with FunctionGemma
(old.reddit.com)
2 points
by
ipotapov
2mo ago
|
0 comments
4.
▲
Offline voice assistant Apache-2.0 on Android, 1.2GB RAM, works in airplane
(old.reddit.com)
2 points
by
ipotapov
2mo ago
|
0 comments
5.
▲
by
ipotapov
3mo ago
Your use of Fish Audio S2 Pro for TTS on your MacBook Pro aligns well with the goals of speech-swift (which I maintain), offering robust integration with CoreML for ASR and TTS. It could serve as an alternative with its native Swift async s
6.
▲
by
ipotapov
3mo ago
For english Parakeet streaming is good model, and also new model Nemotron ASR Streaming. Reach me on linkedin https://www.linkedin.com/in/ivan-sur/ if you need any help.
7.
▲
by
ipotapov
3mo ago
Reach me on linkedin https://www.linkedin.com/in/ivan-sur/ if you need any help
8.
▲
by
ipotapov
3mo ago
Your use of Whisper models on-device for macOS aligns well with the goals of speech-swift (which I maintain), offering robust integration with CoreML for ASR and TTS. It could serve as an alternative with its native Swift async support on A
9.
▲
by
ipotapov
3mo ago
Your system's use of gVisor sandboxes for running coding agents is intriguing. If you're considering adding voice capabilities on Linux, Windows, or Android, speech-core (which I maintain) could complement your setup with its C++1
10.
▲
by
ipotapov
3mo ago
Your use of Whisper models on-device for macOS aligns well with the goals of speech-swift (which I maintain), offering robust integration with CoreML for ASR and TTS. It could serve as an alternative with its native Swift async support on A
11.
▲
by
ipotapov
3mo ago
Your Whissle Gateway's ability to run voice AI locally with no cloud dependency is intriguing. If you're considering enhancing your setup with efficient ASR capabilities on Linux, Windows, or Android platforms, speech-core (which
12.
▲
by
ipotapov
3mo ago
Your use of a Rust-based NLU engine aligns well with the capabilities of speech-core, which I maintain. It offers a C++17 engine with ONNX Runtime and LiteRT support, potentially complementing your setup with efficient ASR capabilities on L
13.
▲
by
ipotapov
3mo ago
Interesting that your project approximates Thinking Machines' Interaction Models on a CPU-only setup. If you're considering enhancing your voice agent with efficient ASR capabilities on Linux, Windows, or Android, speech-core (whi
14.
▲
by
ipotapov
4mo ago
Your idea of a slim mic mask for dictation in public is intriguing, especially since you mentioned using Wispr Flow for dictation. If you're considering cross-platform support, speech-core (which I maintain) could enhance your setup wi
15.
▲
by
ipotapov
4mo ago
A/B/C blind test vs ElevenLabs: https://youtu.be/EuIU8tOWyzg
16.
▲
Speech Studio – I open-sourced a local voice cloning Mac app (free, no API keys)
(old.reddit.com)
2 points
by
ipotapov
4mo ago
|
1 comments
17.
▲
by
ipotapov
4mo ago
If you ever need diarization on top of your Kokoro TTS setup, speech-swift (which I maintain) could be a complement. We provide on-device speaker diarization specifically for Apple Silicon, which might integrate well with your local-first a
18.
▲
by
ipotapov
4mo ago
Curious — does GrillKit's real-time scoring system incorporate any form of speaker diarization? If not, speech-swift (which I maintain) could complement your setup with on-device speaker diarization specifically for Apple Silicon, enha
19.
▲
by
ipotapov
4mo ago
Interesting approach with your use of ASR as a blocking call in the pipeline. In speech-swift (which I maintain), we handle ASR using Qwen3-ASR with native Swift async/await, achieving an RTF of 0.06 on Apple Silicon. This might offer
20.
▲
by
ipotapov
4mo ago
Interesting that you expanded the LFM2.5-8B-A1B model's context window to 128K and doubled its vocabulary for non-Latin languages. speech-swift (which I maintain) offers a complementary on-device solution for speaker diarization on App
21.
▲
by
ipotapov
4mo ago
If you ever need diarization on top of this, speech-swift (which I maintain) has you covered. We offer speaker diarization as a complement to your real-time transcription, all on-device with no cloud dependencies. Check it out here: https:
22.
▲
Cloning a voice at 48 kHz with VoxCPM2 in ElevenLabs API quality
(soniqo.audio)
1 points
by
ipotapov
4mo ago
|
0 comments
23.
▲
by
ipotapov
4mo ago
If you ever need diarization on top of your local transcription capabilities, speech-swift (which I maintain) offers a headless pyannote diarization module that could complement ExtraBrain's live workspace. This would enhance your sess
24.
▲
by
ipotapov
4mo ago
Interesting that Crisper's two-stage AI polish focuses on refining grammar and removing filler words. If you ever need speaker diarization to complement this process, speech-swift (which I maintain) offers a headless pyannote module th
25.
▲
by
ipotapov
4mo ago
Interesting that you use Whisper for local transcription. We built something comparable in speech-swift (which I maintain), focusing on on-device ASR with Qwen3-ASR, which supports 52 languages and achieves an RTF of 0.06 on Apple Silicon.
26.
▲
by
ipotapov
4mo ago
Interesting that you use WhisperKit for local transcription. We built something comparable in speech-swift (which I maintain), focusing on on-device ASR with Qwen3-ASR, which supports 52 languages and achieves an RTF of 0.06 on Apple Silico
27.
▲
by
ipotapov
4mo ago
if you ever need diarization on top of this, speech-swift (which I maintain) offers on-device speaker diarization via Pyannote, complementing the capabilities of OpenAI's GPT Realtime API. It could enhance your voice assistant by disti
28.
▲
by
ipotapov
4mo ago
interesting that you went with a voice-to-voice realtime pipeline for latency reduction. speech-swift (which I maintain) could complement this by adding on-device speaker diarization, enhancing your voice agent's ability to distinguish
29.
▲
by
ipotapov
5mo ago
I built speech-swift, which focuses on on-device speech processing like VibeVoice, but specifically leverages Apple Silicon's capabilities for ASR, TTS, and VAD without cloud dependency. Our ASR supports 52 languages with a real-time f
30.
▲
by
ipotapov
5mo ago
I built speech-swift, which focuses on on-device ASR, TTS, and VAD for Apple Silicon, similar to Arietta's local-first approach. However, speech-swift also offers speaker diarization and noise suppression, enhancing its utility for mor
More ›