GyaanSetu AI

AI、机器学习与 LLM 洞察。

1515 articlesDeep, practical knowledge

Nemotron 3.5 Lightning 上线 AWS,大幅降低企业 LLM 硬件成本

开发人员现在只需点击一下,即可直接从 SageMaker JumpStart 控制台启动 Nemotron 3.5 Lightning,无需配置单独的 GPU 集群、安装驱动程序或使用自定义容器。定价基于 Token 且因地区而异,因此建议团队在投入生产前先验证延迟和准确性。

AI · 2 分钟阅读

研究发现:AI 上下文压缩仅能保留 17% 的用户规则

宾夕法尼亚州立大学的研究人员发现,当大语言模型(LLM)压缩对话历史时,会丢弃诸如“不要使用我的名字”或“更改前先确认”之类的会话约束,导致仅有 17% 的规则得以保留,从而危及安全性。

AI · 3 分钟阅读

谷歌引入前 Relay CEO 助力 Chrome,推动 AI 智能体进入浏览器

Relay 将于 8 月 15 日停止为免费用户提供服务,并于 9 月 14 日停止为付费客户提供服务,从而终止其 AI 自动化服务。与此同时,曾负责 Gmail、日历和聊天产品开发的 Jacob Bank,现已出任 Chrome 产品与开发者关系副总裁,负责领导基于 Gemini 的智能体集成工作。

AI · 5 分钟阅读

Kog Claims 30x Faster LLM Decoding on Existing Nvidia H200 GPUs

Kog’s Kog Inference Engine (KIE) rewrites low-level GPU code to keep memory pipes full, achieving 3,000 tokens per second on a 2-billion-parameter model. Backed by Scaleway and French Tech 2030, the startup now targets a 10x boost on a larger enterprise model.

AI · 5 分钟阅读