Back to feed
AI in Industryr/LocalLLaMA · June 25, 2026 · 4w ago

LFM2.5 230M running in-browser at 1,400 tok/s using custom WebGPU kernels

<table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1ufii9b/lfm25_230m_running_inbrowser_at_1400_toks_using/"> <img src="https://external-preview.redd.it/ZzBzdGIwM3R5ZzloMbNWdyfcno-xz1w51HUC8uIZ38Kd2Cw6Gsvbd8-SkCU1.png?width=640&amp;crop=smart&amp;auto=webp&amp;s=f7aeb7ccb5763d423154bc64999c193a0b9532bd" alt="LFM2.5 230M running in-browser at 1,400 tok/s using custom WebGPU kernels" title="LFM2.5 230M running in-browser at 1,400 tok/s using custom WebGPU kernels" /> </a> </td>
Open original

The Daily Drop

Join 1,000+ people who read this first.

Related stories

Anthropic Resets Rate Limits for All Claude Users

Anthropic has reset both 5-hour and weekly rate limits across all Claude user tiers, effectively giving everyone a fresh quota allocation. This appears to be a system-wide refresh rather than a permanent policy change. Why it matters: - Teams hitting rate caps during high-usage periods (end of sprint, campaign launches) get immediate relief without needing to upgrade or wait for natural reset cycles. - The move suggests Anthropic is actively monitoring capacity constraints and willing to manually intervene when bottlenecks emerge, a positive signal for enterprise reliability. - For developers building agentic workflows or batch processing tools, this confirms rate limits remain a real constraint to design around, not just theoretical guardrails. - Marketing teams running large content generation batches should anticipate similar future resets if infrastructure keeps pace with demand growth. This likely reflects either a capacity expansion or a response to user feedback about quota friction. Watch whether Anthropic announces permanent tier adjustments or keeps using ad-hoc resets.

X / @claudedev · 1w ago