The subsidized subscription model won't last, API pricing "feels" closer to a true sustainable business model.
It’s so cost effective I can offer a generous free tier since my goal isn’t to make money with it.
I'm on a team that develops a Bible study app, and we're all relatively content with how the basic models converse regarding scripture. Even as far back as GPT-4 was excellent. They occasionally have minor hallucinations (a dealbreaker for a production app), but they do an excellent job with theology and Bible scholarship, given reasonable guardrails.
I'll admit I'm coming from the perspective of "should we be implementing this?" It seems, on the surface, that a strong embedding-based verse retrieval covers the bases at a microfraction of the cost.
If you're interested, check out the development server where we're working on this. You navigate to the search (magnifying glass) and then hit "Meaning". Sorry for the confusing route; we're still deciding on back-end details and haven't focused on the front yet.
What's difficult and doesn't have to be with philosophy/ spirituality is to find relevant bits off situation, theme etc.
This app does that very well, LLMs are good at entity recognition.
One feature of the app is that all scripture is verified and what’s show to the user doesn’t come from the LLM at all and instead a trusted source.
I think exploring scripture this way does not alleviate you from struggling to learn and apply it. It hasn’t for me.
But there are ways to control and constrain the LLMs and what the user is presented with.
These are all top of mind for me and why I felt there could be a better option than asking ChatGPT directly.
Deepseek v4 Pro prices with Opus 5 perf would be freaking unbelievable!!
This is probably a dream.
https://artificialanalysis.ai/models/deepseek-v4-flash?intel...
First, your direct comparison, Deepseek V4 Flash 0731 (max effort) $0.03 (rounded up) per task @ index 50.
OpenAI Luna:
* high effort $0.03 (rounded down) @ index 46
* xhigh effort $0.04 @ index 49
* max effort $0.07 @ index 51
So I would say a fair statement would be "OpenAI Luna between 2x and 3x the price of Deepseek Flash, what you get is 2 to 5 times faster inference"
The cheapest OpenAI model that beats it is OpenAI Luna (max effort) $0.07 @ index 51 (if you take the rounding out it summarizes to triple the price for similar performance), but still close to 3x faster.
And can SOMEONE please tell artificialanalysis that using dark blue for both Deepseek AND OpenAI is an especially unfortunate choice of colors, especially today?
For simple tasks, they're already saturated, and you'd prefer the faster model, so that you can have a realtime/interactive-ish experience.
Or to put it bluntly, it's cheaper if you don't value your time. That goes for smaller models in general -- need more handholding, more correcting -- but the Chinese ones are slower on top of that.
As for speed, Sol on Low is faster than Luna on most settings.
It’s also so inefficient, when they release the full performance numbers it’s not going to be good.
One example, it takes about 3.6x more tokens to finish the same work as Gemini Flash 3.6.
[0] https://petergpt.github.io/bullshit-benchmark/viewer/index.v...
Similar price? Doesn't make sense. Maybe they meant power, capability or speed?
Am I just using it on tasks that makes it go on forever vs these benchmarks that are short&sweet, or something like that? I've been throwing bunch of identical prompts at different models at the same time, and when comparing hy3 and K3 I've never once had K3 reason less than hy3, as just one anecdotal data point.
The specific agent is focused on getting precise and on point answers about a codebase.
The starting point was nowhere near. E.g. asked why was X implemented in a certain way it would give bogus answers when the real answer was that there was no reason at all.
The benchmark included more than 50 questions or different difficulty.
But when the agent was improved in its prompt and rooting it was impossible to have it perform worse than closed source sota.
Just to say that the quality of the harness is as important as agents intelligence.
Or a benchmark to benchmark benchmarks?
benchmark website benchmark is indeed a benchmark that benchmarks websites with benchmarks (but it can be shown outside websites as well, it's not picky)
The ban on these open models is coming within weeks, if not days. As usual, the excuse will be "national security".
For example: no government contract to any company who uses even one vendor in it's entire chain of dependencies, who uses such open models.
They can extend this further by laying more conditions, such as: any company dealing in this-this field can only use models "officially" approved as "safe". Rest you can guess how easy it would be to get that "safe" rating for such open models.
I claim the CCP will wise up within 2 years, possibly much much sooner, and ban their own companies from open sourcing to prevent the Americans from acquiring the capabilities.
Despite all the nonsense claims of China distilling US models, the reality is that the Americans absolutely do distill these free Chinese models, and distillation when full logprobs are available (i.e. you have access to the weights of the model) is an order of magnitude better than when you don't.
Yes, Chinese open weight models in the short term harm US closed source model providers bottom line. In the slightly longer term, "showing your hand" and publishing both the architecture innovations and the models weights will be too dangerous for the CCP to allow. This is triply true if they can release a model that beats the Americans on most benchmarks.
I've already warned investors that this is probably the closest open weight models will ever get to closed access.
https://www.businessinsider.com/xi-jinping-open-source-ai-us...
People on HN downvote objectively correct information because they don't like it 24/7. There's a reason the creator of Zig left and gave the computer version of a middle finger on the way out to HN!
commenting about voting is also something the HN guidelines warns against:
> Please don't comment about the voting on comments. It never does any good, and it makes boring reading.
Daily reminder that improving your samplers from the garbage default top_p/top_k to min_p or subsequent methods dramatically improves the performance of these models, and makes most quantities like measured "verbosity" and subsequent calculations of "intelligence per token" meaningless
Daily reminder that no one, including within academic AI research, AI engineers, normies, etc takes LLM sampling seriously enough.
I have to admit it rarely comes up in the coding tasks I usually give to LLMs.
But you already know that.
Any normal user is much more likely to ask questions to which the Anthropic and OpenAI models do not answer, than to ask questions about the modern Chinese history, to which a Chinese LLM will not answer.
If you poke it just a few times, however, you get to the point where it will eventually say (paraphrasing) that basically only Israel, the US state department, and the ICJ say it's not a genocide.
That is to say that it's framing it as some sort of tricky complex question when it's not. And when interrogated, it basically admits that the only people who dispute it are Israel and it's supporters.
The majority of the world is religious - doesn’t mean the debate on religion isn’t a complex question.
The majority of the world approved of slavery historically.
The majority of countries have ethnically cleansed their Jews, many of them in living memory.
When interrogated you will find that the only ones asserting the war in Gaza is a genocide are people who were anti-Israel anyway.
Plus a size you can genuinely run at home (36GB vram, 200gb RAM).