Top
Best
New

Posted by willwhitedc 23 hours ago

Desert Ant Labs: local, fast models that run on device(desertant.com)
448 points | 96 commentspage 2
library8848 22 hours ago|
Shiny layer of marketing and proprietary code on top of open models?

Voz is Parakeet 0.6B v3

Clear is DeepFilterNet 3

Ear is the language predictor from whisper-tiny

...

sudb 19 hours ago||
It is now also reasonably straightforward if you have access to frontier LLMs, a recent-ish mac and a recent-ish iPhone to point them at the job of porting a given model to run on the ANE - it's a reasonably easy task to hill-climb at this point!
Muromec 20 hours ago|||
AI beige theme and obnoxious AI writing signaled as much.
joshuat 18 hours ago|||
but they're European!
sipjca 21 hours ago||
seems to be…
ricardobeat 17 hours ago||
> Ranks a transcript's best non-overlapping moments: each clip gets scored and ranked. Build strong selections or unique editing features to pick the best sentences in video or audio recordings.

This is very impressive for a 248MB model. I wonder how good the results are, as an LLM 10x the size is still quite bad at that.

markdog12 21 hours ago||
> opinionated on-device intelligence

> Hate speech triage. On-device moderation that flags hateful, abusive and threatening text

What could go wrong here?

lemome 21 hours ago||
I don't think you understood. It means these are specialized models. Their toxic model could be ideal for video game lobbies without investing a ton of money if you're an indie dev

This also could be ideal if you want your child to play online to have auto-censorship

ch_sm 19 hours ago||
I don’t know if i want to expose my child to auto-censorship (or online gaming anyway).
Muromec 20 hours ago||
1. Not having human in the loop to review it because humans are expensive.

2. Having human in the loop to review it and subject said human to the worst other humans produce.

lukevp 19 hours ago||
I would love to use Voz and Ear, but I’d need a version that is competitive with other audio transcription LLMs for platform availability - meaning macOS, Windows and Linux, and supporting GPUs if available.
nullbio 22 hours ago||
This is a cool idea. The most useful one for me would be something that can process pdf files into a json schema. Title and tag generation from a post would also be useful. I'm interested in web app though.
pveugen 21 hours ago||
OCR on steroids. Our Schemer model will soon be available (free form text to structured JSON). Once that lands, we want to jump into image to JSON.
capevace 19 hours ago||
But Schemer will be text-only, correct?

Once image support drops this could be a cool addition to https://struktur.sh. Will it be on OpenRouter too or will all inference have to be self-managed? I’m thinking about server-side use cases, where budgets are low, so small models shine.

edit: ah I saw mainly iOS for now. But an integration would be possible on macOS then, right?

pveugen 17 hours ago||
Schemer is text only. But we're working on sort like models for images to data too.

We aim to make all our models available cross-platform. There's a subset currently only iOS/macOS, because they take some more effort to port over to Android and web.

Check out the CLI for easy experimentation: https://github.com/Desert-Ant-Labs/desert-ant-cli#desert-ant...

ironsmoke 16 hours ago||
You should check out IBM's Docling.
faangguyindia 19 hours ago||
Cool! is there a local model for LLM command approval?
illright 22 hours ago||
I wonder why they only support Apple platforms, citing CoreML. Doesn't Android have a similar framework, ML Kit?
pveugen 21 hours ago|
We plan to make most our models available for Android and web too. Some are a bit harder to port to the different platforms and will take a bit longer to properly land on Android or web. Mostly sequencing (Voz, Clips, Title). Soon!
Dwedit 16 hours ago||
Testing out Tongue, it detected "馬鹿外人" as Chinese.
eproxus 13 hours ago|
It does say this: "Certain by script: these characters belong to only one language, so the model never ran."

So I think it's some buggy code in their UI that even prevents the model from running.

viccis 18 hours ago||
I wish we could pop a tiny one into my phone so that when I type "Will see you" and swipe the word "later" it chooses that instead of "lasso"
agcat 18 hours ago|
Honestly this is a really cool idea and i am glad that there are companies being built in this space. This is closest to the vision of what i want to do next.
More comments...