5h ago 5小时前
How I Taught Claude Code to Offload Grunt Work to a Local 4B Model 我是如何教会 Claude Code 将琐事卸载到本地 4B 模型的
Claude Code subagents cannot route to different models—they share the same base URL, making model isolation impossible through subagent configuration ... Claude Code无法通过subagents或环境变量切换模型,正确方案是使用Skills + Bash组合将本地小模型作为工具调用
通过Python脚本桥接llama.cpp的Anthropic兼容API,实现Claude对本地Qwen3.5-4B的按需委托
遇到"reasoning trap...
Claude Claude LLM 大模型 Agent Agent Code Generation 代码生成 GPU GPU
9h ago 9小时前
TAI #219: AI, Cancer and the Future of Personalized Medicine TAI #219:AI、癌症与个性化医疗的未来
Anthropic's AI agents successfully coordinated protein design workflows, achieving a 26.8% success rate across 1,320 designs and producing at least on... Anthropic AI代理在蛋白质设计中实现26.8%成功率,15个靶点中14个至少产生一个成功蛋白
Moderna与Merck个性化mRNA癌症疫苗INTerpath-001 III期试验达标,1,137名高危黑色素瘤患者参与
NVIDIA AVO Agent基于Claude Opus 5在AR...
Healthcare AI 医疗AI Agent Agent Multimodal 多模态 LLM 大模型 Claude Claude
10h ago 10小时前
From Transaction Graph to Replayable Audit Trail: Fraud Scoring on TuringDB 从交易图谱到可回放审计轨迹:TuringDB上的欺诈评分
Structural fraud detection leverages AI to identify complex, multi-layered fraudulent schemes that traditional rule-based systems miss
Laundering-patt... 结构性欺诈检测利用人工智能识别传统基于规则的系统所遗漏的复杂、多层欺诈方案
洗钱模式分类利用向量搜索以高精度映射和分类金融犯罪模式
基于提交历史构建的可重放审计轨迹提供透明、可验证的检测逻辑和决策记录
这三个组件的整合创建了一个全面的、可审计的AI驱动欺诈检测流水线
Finance AI 金融AI Security 安全 Embedding Model 嵌入模型 Research 科学研究
11h ago 11小时前
Why vLLM and SGLang Are Replacing Ollama for Agentic Workflows 为什么 vLLM 和 SGLang 正在取代 Ollama 用于智能体工作流
Agent loops repeatedly re-send nearly identical prompts across hundreds of steps, creating significant redundant computation
vLLM and SGLang implement... Agent循环场景中会重复发送几乎相同的prompt数百次,造成大量重复计算
vLLM和SGLang框架在步骤间保持prompt缓存,有效减少重复推理开销
Ollama在五次请求后丢弃缓存,可能导致长Agent循环性能下降
LLM 大模型 Inference 推理 Agent Agent Open Source 开源 Deployment 部署
12h ago 12小时前
How Aiden Agents Survive Running Out of Context Mid-Task: A Technical Deep Dive Aiden 智能体如何在任务中途耗尽上下文时生存:技术深度解析
Context window overflow is a structural certainty for multi-step AI agents, occurring either through cumulative history growth or single oversized too... AI agent在长任务执行中必然遭遇上下文窗口溢出,分为累积溢出(多轮对话超限)和单次事件溢出(单次工具调用返回数据过大)两种类型
Aiden实现了三层协调机制:上下文压缩(有损)、会话切换、以及基于外部持久化文件的恢复,通过读取Continuation ID、保存状态、保存错误和输出尾部四个证据...
Agent Agent LLM 大模型 Research 科学研究 Deployment 部署
13h ago 13小时前
The Index Is Not the World 索引并非世界
A language model reading the Epistolæ corpus of medieval women's letters recovered 203 named individuals across 30 letters, compared to only 54 people... 传统基于索引的网络图会严重低估历史社交网络规模:在30封中世纪信件样本中,元数据图仅连接54人,而阅读内容后发现实际涉及167个真实人物
AI阅读能够恢复被网络图折叠的信息:包括被简化的群体(如55名骑士被压缩为单一节点)和从未出现在索引中的隐藏人物(占三分之二)
信件中的"代求"行为(interc...
LLM 大模型 Research 科学研究 Dataset 数据集
14h ago 14小时前
Dissecting llama.cpp, Part 1: From GGML to GGUF, and Why llama.cpp Is Just the Wrapper 剖析 llama.cpp(第一部分):从 GGML 到 GGUF,以及为什么 llama.cpp 只是一个封装层
GGML is a lightweight C/C++ tensor library and model serialization format designed specifically for running LLaMA on consumer CPU hardware without GPU... GGML是llama.cpp的底层数学库与模型存储格式,专为消费级CPU推理设计,支持激进量化与零依赖部署
GGML采用图执行而非PyTorch的急切执行模式,通过预构建计算图实现一次性内存分配,显著提升批量为1的自回归生成性能
GGML文件格式历经GGML→GGMF→GGJT三阶段演进,GGJT通...
LLaMA LLaMA Open Source 开源 LLM 大模型 Inference 推理 Quantization 量化
15h ago 15小时前
Understanding the Impact of AI on Job Markets 理解AI对就业市场的影响
AI transforms jobs through five distinct mechanisms: displacement of routine tasks, augmentation of knowledge work, creation of new occupations, compr... AI对就业市场的影响是多元复合的,分为替代(Displacer)、增强(Augmenter)、创造(Creator)、均衡(Equalizer)四种力量同时作用,而非简单的"取代人类"叙事
WEF《2025年未来就业报告》预测:到2030年将有9200万个岗位被替代,同时创造1.7亿个新岗位,净增7...
Research 科学研究 Ethics 伦理 Policy 政策
15h ago 15小时前
Why NVIDIA's Open Weight Nemotron 3.5 Lightning Shows Their Strategy To Remain The AI King 为何NVIDIA开源Nemotron 3.5 Lightning彰显其保持AI霸主地位的战略
NVIDIA reported $75.2 billion in Data Center revenue for its latest completed quarter
Data Center segment accounts for approximately 92% of NVIDIA's t... 英伟达最新完整季度数据中心业务收入达752亿美元
数据中心业务约占英伟达816亿美元总收入的92%
这凸显了英伟达在AI基础设施和GPU驱动计算领域的主导地位
这些数字凸显了企业AI采用和GPU需求的巨大规模
LLM 大模型 Open Source 开源 GPU GPU Inference 推理 Research 科学研究
16h ago 16小时前
How Artificial Intelligence Detects Supply Chain Risks Before They Become Problems 人工智能如何在问题爆发前检测供应链风险
AI is transforming supply chain management from reactive problem-solving to proactive risk prevention by continuously analyzing thousands of interconn... AI通过持续分析数千个互联信号,实现从被动响应到主动预防的供应链风险管理范式转变
机器学习能够发现人类无法察觉的隐藏模式,如特定天气条件下的港口延误、供应商绩效渐变等早期预警信号
图机器学习可映射供应商、工厂、仓库、物流商和零售商的网络关系,揭示隐藏依赖关系并预测连锁影响
数字孪生技术允许企业在虚拟...
Research 科学研究 Deployment 部署
16h ago 16小时前
V-JEPA 2 Is a World Model. Just Not the Part Everyone Means. V-JEPA 2 是一个世界模型,只是不是大家所指的那个部分
A humanoid robot achieved a 100-meter sprint time of 9.39 seconds at a competition in Beijing, surpassing Usain Bolt's human world record of 9.58 seco... 一台人形机器人在北京举行的比赛中以9.39秒的成绩完成100米冲刺,超越了尤塞恩·博尔特保持的9.58秒的人类世界纪录。
这一里程碑标志着人形机器人在动态移动和实时平衡控制方面取得了重大突破。
该成就展示了近年来具身人工智能和机器人硬件能力的快速进步。
比赛形式可能同时测试了速度和稳定性,凸显了双足...
Robotics 机器人 Research 科学研究 Multimodal 多模态 Training 训练
17h ago 17小时前
Learn These 5 AI Terms and You'll Understand More Than Most People Who Use AI Every Day 学会这5个AI术语,你的理解将超过大多数每天使用AI的人
AI operates on tokens rather than words, which directly impacts pricing, processing speed, and memory limits
The context window determines how much te... 理解AI的5个核心概念:Tokens、上下文窗口、Temperature、幻觉和RAG,无需编程背景即可掌握
Tokens是AI处理文本的基本单位,直接影响成本、上下文长度和响应速度
上下文窗口限制AI同时"看到"的内容量,超出范围会导致早期信息丢失
Temperature参数控制AI输出的创造性与...
LLM 大模型 Training 训练 Inference 推理 Programming 编程