Ember-1
Discussion 244 comments
this is our preferred open weight token vendor
this work may explain why recent models like qwen-3.8-flash and MiMo-2.6-* have not made it into their offering, which has given me reason to pause my excitement for Fireworks
…trust me bro.
It’s obviously easier to believe when they’re not training models.
Eh, anyway this whole thing is just an ad:
> Looking to take Ember-1 one step further, and optimize it for your use case? We are also launching training support for Ember-1, enabling enterprises to build customized, token-efficient models tailored to their needs with their own data. The future of open models is specialized models trained on your specific workload.
Probably, I guess, fancy serverless infrastructure actually makes virtually no difference to hosting really large models that people want to use, and “just” being an inference provider for open weight models turns out to have no moat.
So this is a bit of a pivot to “use our training infrastructure too…!” imo.
Pivot? Sure. Go them. Not what I signed up for though. /shrug
Hong Kong I think used to be a popular option as well, before the handover from the UK.
Being terminally 'online' can make people cynical about everything, but have to temper things with reality a bit too.
They've been using the Pelican test in the /r/Codex subreddit with some success. One major finding is that OpenAI has been silently degrading the model while charging Astra prices. Another finding is that even when the model has not been silently degraded, pelican quality is significantly lower. Sometimes comically so. The general consensus right now is that the new GPT-6 Sol model is an updated Terra model. Many intelligence metrics are roughly similar. Meaning the most recent model updates were an attempt to rebalance compute rather than improve intelligence.
Ultimately I've never seen users as upset about GPT-6 Sol/Luna than I have right now. Even Astra has been noticeably degraded for me and everyone else I have asked. This is compounded by the fact that Opus 5.5 is a generational improvement at an affordable price. There is currently no competition.
https://arxiv.org/pdf/2307.09009
the accusations are quite a few because people notice.
—"Benchmarks!"
...I'll tell that they can be gamed so easily, and they are on a consistent basis.
With Sol 6 I am back in a world where the model writes bad code because it is lazy ("You're absolutely right, I did not [do it properly] because I did not want to edit [a normal amount of files]").
Ember isn't picked yet. In planning, Opus 5.5 wins under the planning weights. In code, GPT-6 Sol dominates it: also 10/10, but with a higher quality score and a lower estimated cost. Ember has no intelligence index, so its starting score is only 0.73, which holds its 10/10 down to 0.954 against Sol's 0.975.
[1] https://philippdubach.com/posts/jev-model-router-for-pi/
6 or 5.6? Because 6 is hot garbage
Kimi K3 itself isn't FOSS. Speaking of reciprocity: Fireworks is presumably paying Moonshot serious money for the right to do what they are doing here, since Kimi's license[0] excludes commercial inference providers (such as Fireworks) from gratis use. It requires them to: "...enter into a separate agreement with Moonshot AI before using the Software or its derivative works..."
[0] https://huggingface.co/moonshotai/Kimi-K3/blob/main/LICENSE#...
I suspect the advantage that catapulted Linux ahead of the establishment was less technical potential and talent and more organizational advantage. That's not to diminish the technical talent of the Linux crew, but them being unencumbered gave them more degrees of freedom. The rest is history.
So as long as the AI companies don't succumb to "big company" dynamics, they can outlead. To wit: Open AI and Anthropic are kicking Google's ass.
Google might not have compelling frontier offerings, their chat harness is complete garbage compared to any other lab (in large part due to a bizarrely badly designed harness where something like code execution requires the prompt to undergo some sort of classification step, no idea what they are doing).
But they absolutely kill in terms of usage offerings. Google lets one subscription be used by *SIX* different google accounts on a family plan.
Plus I currently literally get *$40/month* of Gemini API credits on developer.google.com because they gave me a $10/month grant 4 times.
They give you 200 cloud compute units on google collab, this literally lets you spin up an H100 for around 40 hrs or something if you want to try spinning up local models.
You get Jules (huge allotment btw), Image gen, Video gen, Music Gen, antigravity usage, 5 TB of cloud storage, Notebook LLM...
in one month, Google actually went cash-negative. [0] even still, they are subsidizing their stuff a lot less, have the most opaque and variable limits, and increase adoption through bundling and shuffling features. I can't even share my Google One storage without subscribing to a Google AI plan anymore, but previously any plan except Google One Lite was shareable.
if you tell me that's not enough to go after frontier, then how much money are Anthropic and OpenAI burning?
[0]: https://www.techspot.com/news/113214-google-records-first-ne...
Indeed. And when you have freedom to play, you are able to find new stepping stones that you didn't anticipate. And you can combine stepping stones in new ways to make new discoveries.
Greatness cannot be planned.
The lock-in is less pronounced as it is with AWS or MS.
I can also see the argument for providing a post-training service from a customer acquisition perspective: "hey, we can fine-tune this open weights model, so it both gives better/more predictable results than OpenAI/Anthropic and also is cheaper. And btw, once we've won your business, please run this model on our infra."
But what I'm struggling to understand is fireworks spending a bunch of money (on salaries and compute) releasing a frontier model that is going to rapidly fall behind the frontier. Is this "just" advertising for them, both for customers and also for hiring? Or are they actually trying to stay on the frontier? If so, to what end?
It's unclear if they can do this systematically and it's unclear if they can do it better than others. But, lots of things are unclear in AI at the moment, this doesn't seem outrageous on the surface. And, it could just be marketing. And it could be the first option with the backup of the second.
I'd expect that their business strategy is to compete in more markets, and if successful, they can capture more value. This is the "easiest" for them as they already have GPUs, a training environment etc. For that platform it's not the worst if there's an internal customer team that can help shape the future and provide immediate feedback, and if it results in a good model, even better.
Other things I'd not be surprised they offer in the future in the same vein: A multi-model harness, coding agent (cloud and local), and maybe at a later point in time even a CPU-only cloud compute product.
here, it does less work, it's more like driving half as far but still paying the same total cost
this being said, K3 and E1 models are priced the same at $3.00 / $0.30 / $15.00
Wouldn't the op be more correct with their gas/distance comparison?
Because both cars get to the same end destination (complete the same task).
Unless your end goal is to see the token numbers go up, but I'm not sure why that would be of interest.
Will be taking Ember-1 for a spin on Monday and hopefully enjoy those better MPGs
https://aibenchy.com/compare/fireworks-ember-1-high/openai-g...
however the listing on open router has a `/fp4` suffix, so perhaps this is an unlisted, quanted model for a lower price?
we require ZDR and Fireworks provides that on contract, so for us they are the same price
industry standard is price-per-1M tokens, don't do something different, even Google caved and moved from their char based pricing to tokens (the fundamental unit of computation in ai)
GPUs are rented in $/h, like every other piece of hardware in cloud
This is partly the appeal of Jev et al; having a quick model for simple tasks, that doesn’t require that much thinking
It’s amazing all the workflows that models like that can unlock. And yes, classifiers and other ML models have been around for a while for these types of tasks, but Jev has made it easy and cheap to play and experiment. This in turn, is incentivizing people to try them for a bunch of stuff, unlocking creativity and producing a lot of new cool (and eventually potentially very useful) applications
1. Evals (once you have your rubric defined and tuned using a reasoning model, jev can be great for running periodic evals especially those that run daily.
2. e-commerce catalog classification 3. quick search using anything as context and query mapping to a pre-defined set.
For example, a typical/stock LLM can’t really play Doom in real time, but a Jev-like model can. Just because of latency
Of course, if you want the best Doom player, there are way better and faster adhoc models
Analysis paralysis stifles not just human intelligence, but other intelligences too.
The more options you have, the harder it becomes to be satisfied with the one you picked.
> task and environment feedback
> on-policy planning and learning
> feedback connects decisions to their consequences
These are deliberately the least informative phrases you could possibly use to describe what you have done, while still being in the realm of words that go over a generic investor who has no idea whats going on and may be dazzled by sciencey sounding language.
Cursor compose 2.5 article where they used and described on policy self distilation was actual alpha.
Not suggesting this is right or wrong, but is sort of the nature of the technology.
The words of a license are what the license is.
The pareto frontier needs clearer distinction. Benchmarks miss half the story. What, if any, capability is lost by the token reduction (for example, was it like super awesome at Golang before and now kind of sucks? that kind of distinction).
Looking at the hamsters drawings in this comparison, never saw any other model make such a similar version:
https://aibenchy.com/compare/fireworks-ember-1-low/google-ge...
I'd be interested to hear what people found most interesting about Ember, and what kinds of follow-up research or educational material would be useful to you all?
I see that with Opus 5, it started thinking like crazy in the last few days , I don't think my workflow is that complicated, still it gets into thinking mode and stays there
It's pretty important to understand if your own work domain is one where the last 5% matters. In a lot of day-to-day software engineering tasks, it doesn't, and one can get crazy mileage out of the cheaper models. OTOH, if you are performing novel research, that last 5% may be worth whatever it costs...
On the other hand, if it's cheap to tell whether you got a good solution, and you think the 90 and 95% apply to your task blend, then it's almost always worth trying the cheap model first.
> creative
Choose one.
That's just vibes, though.
Also, did I miss a memo? Suddenly every article on AI seems to be talking about the Pareto frontier - or have I just not been paying attention?
Kimi K3 with less reasoning tokens isn't exactly exciting either, and particularly so if the license is less open than original Kimi K3.
codeAnyone knows more or used this model?
"Pareto": 8 hits
"Opus 5.5": zero hits
A tiny, smaller than 1b parameter model, fine-tuned, can kick ass for constrained work. I do not have a lot of budget, I fine-tune only on a 16GB M4 Mac Mini. But that also tells me the potential is wild. Progress has been slow since I moonlight on this.
I have been trying to build a set of models + agents for full-stack development, where each model does only a small piece, like take user prompt and break into backend/frontend tasks. Then a Rust+Diesel model, a Rust+Auxum model, a Solid+Router model and so on. I know this is wild but this is just theory - can 5 or 6 Qwen 3.5 0.8b models do full-stack web development? My hunch says they can, better than what most people expect. Heck, with a good harness, it might beat all the cheaper models for the specific task, like Haiku or Luna.
That being said it’s much easier at the moment to continue to use the frontier providers for most general tasks, that is the argument I’ve heard.
For creating these types of fine tuned local models, on constrained hardware for inference, I do think this is the way to go for specific tasks too!
Have you actually fine-tuned yourself? Email categorization comes to mind and there are tons of non-LLM approaches even that will give fantastic results. How did spam filters work before LLM?
I think LLMs just made us think that is the only way. It is not.
I have a single command that fires up llama.cpp on cpu only using gemma4 e2b, answers a single question from the command line and exits. This takes about 3 seconds to load from an SSD, and is smart enough to solve exactly these "remind me of the syntax" scenarios if you dont wanna switch to a browser.
Or are the subagents generating your training data using a closed/paid model?
For everyday work that happens frequently it's better to have a tiny specialized model instead of making billable API calls or turning your laptop into an 80W space heater for 20 seconds to run a general purpose model.
The large models can be used to generate synthetic training data. Tell them to make up 100,000 tasks paired with the resulting output as a 1-time cost. Then use that to train a small model.
Think of it as distillation, but focused on a specific task.
Which 5$ is a pretty easy sell if its useful in any way, It's pretty easy to justify a purchase if its yours forever and doesn't use much CPU so is easy to run I mean people were spending 1000$+ on mac mini setups to run local llms or run remote agents.
This particular example is maybe a niche, but 1400 people can use a few hundred queries in a reasonable amount of time.
I think OP's point remains, if you generate 140k pairs, your local model would need to run that many to offset having just used the generator (SOTA or not) model to begin with.
I wonder if another approach if latency is a concern is just to do a two shot pass with Jev (perhaps given small context you'd want one to match command, then one to match args of given command) would be an extremely fast, and cheap way to do it - rather than training your own.
I ask cause would this be a kind of model distillation?
I have a small model I'm looking to train on some data, and I have some real live data but I'd love to be able to extend it.
edit: others have asked any you have replied "soon (tm)", looking forward for that day
Not my project
I am building a natural language to CSV/Excel commands for a "wrangler" type desktop app. Same issues. The MVP is being built with parsers of sorts, entirely code generated. Then I want to fine-tune a tiny model at some point.
https://github.com/brainless/baho
we dont know what the result is and how its impressive.