2026-08-07 07:27:16
https://www.mayerowitz.io/blog/mario-meets-pareto
这篇文章探讨了《马力欧卡丁车 8》中如何选择最佳赛车配置的问题。游戏中有驾驶员、车身、轮胎和滑翔翼四个部件,每个部件都有多种选择,组合数量惊人。
作者引入了一个多世纪前经济学家帕累托提出的“帕累托前沿”概念。通过比较速度和加速两个关键属性,可以筛选出那些“不被支配”的选项,即不存在另一个选项在两项属性上都优于它。例如,酷栗宝在速度和加速上分别被凯瑟琳和桃花宝宝压制,因此是低效选择。
帕累托前沿能帮你客观地剔除次优选项,但最终选择仍取决于你的个人偏好和游戏风格。文章最后指出,这种多目标优化问题在生活中也很常见,比如选择便宜又美味的餐食、高薪且轻松的工作等。当你的偏好权重不确定时,帕累托前沿能帮你缩小选择范围,专注于高效选项。
https://news.ycombinator.com/item?id=49195231
https://www.crimepaysbutbotanydoesnt.com/reading-list
这是一个关于如何自学植物学的网页指南。
内容主要分为几个部分:
学习心态:鼓励初学者不要被专业术语吓倒,利用互联网资源主动查询不懂的词汇和概念。
核心概念:解释了使用拉丁学名(如 Cedrus)的重要性,因为通用名(如“雪松”)容易混淆。同时介绍了现代植物分类学是基于进化关系(单系群)来划分的。
推荐书籍:列出了几本关键教材和读物,包括:
实践操作:简要介绍了如何制作植物标本(压花),例如将枝条压在旧书或速写本中。
https://news.ycombinator.com/item?id=49192566
这是一个关于 DeltaDB 版本控制系统的产品介绍页面。
DeltaDB 是一种在代码提交之间记录工作过程的版本控制工具,能将每一次变更与产生它的对话关联起来。
主要功能包括:
页面还提供早期访问申请入口,需填写邮箱和 GitHub 用户名。
https://news.ycombinator.com/item?id=49187256
https://runarcn.no/android-to-linux/
作者因不满谷歌对安卓开源项目的控制,如依赖谷歌服务、限制自定义 ROM 开发、强制 AI 功能及限制应用安装,决定将手机从安卓系统切换至 Linux。作者拥有 Fairphone 4,最终选择了 SailfishOS 系统,该系统基于手势导航,应用框架美观,且可通过 SSH 远程控制。但存在 Python 和 glibc 版本过旧、Waydroid 和 GPS 故障等问题,部分社区应用质量不佳。Ubuntu Touch 虽支持 Waydroid,但通知和剪贴板不同步,原生应用体验较差。作者仍需依赖安卓设备访问银行、政府服务及 Uber 等应用,计划使用备用手机 Galaxy A17 通过热点完成这些任务。作者将记录使用体验,并可能购买 Jolla Phone 2。
https://news.ycombinator.com/item?id=49188022
https://neon.com/blog/how-castform-neon-beats-frontier-models-on-price-and-efficiency
Neon 与 Castform 合作,推出了一种成本极低的检索模型训练方案。核心思路是利用 Neon 的 Lakebase Postgres 数据库和搜索扩展,结合 Castform 的强化学习后训练技术,将企业数据库中已有的原始数据转化为高性能的 AI 搜索模型。
文章指出,优秀的 AI 代理需要强大的上下文检索能力和模型决策能力。传统依赖嵌入搜索和 GPT-5.6 等前沿模型的多轮搜索成本高、速度慢。而 Castform 通过后训练,能让成本低 100 倍的开源模型在特定搜索任务上达到甚至超越前沿模型。
Castform 的训练流程完全基于 Neon:使用 Lakebase Search 进行数据存储、合成训练数据、执行训练中的搜索工具调用以及最终推理。企业无需准备复杂的训练数据集,Castform 能自动将内部文档、产品记录等数据转化为问答任务,并管理强化学习循环。训练过程中,Neon 的动态计算扩展能力能有效应对高并发的搜索请求,而 Neon 的分支功能则为有状态代理的训练提供了隔离的环境。
https://news.ycombinator.com/item?id=49186762
https://randsinrepose.com/archives/blade-runner-title-cards/
这篇文章是一篇关于字体排印学的博客文章,作者以《银翼杀手》的片头字幕为切入点,探讨了字体设计的功能性与情感表达之间的悖论。
文章首先指出,优秀字体设计的最高境界是让读者忽略字体本身,专注于文字内容,但字体总会不可避免地传递出某种“感觉”。作者以自己日常使用的编程字体为例,列举了六种备选方案,说明即使是追求功能性的等宽字体,也会带来不同的个人感受。
随后,文章将焦点转向《银翼杀手》的片头字幕,揭示其精妙之处:整个片头仅使用了一种名为“Goudy Oldstyle”的字体,但通过全大写、小型大写、斜体、红色强调以及类似书籍排版的首行缩进等手法,营造出强烈的未来主义氛围。作者特别指出,片头中“Replicant”一词首次出现时使用了红色斜体,与电影标题颜色一致,暗示了影片核心冲突。
文章还对比了《银翼杀手》工作版片头使用的“Impact”字体,认为那是缺乏思考的平庸选择,而最终版对细节的执着正是其“令人惊叹”的原因。作者由此引申到产品设计领域,强调那些看似微不足道的细节决策,最终会汇聚成用户能感受到的“集体声音”,决定一个产品是平庸还是卓越。
https://news.ycombinator.com/item?id=49189287
https://blog.fogus.me/llm/born-against.html
这篇博客文章探讨了为什么爱好编程社区(如 OSDev、LangDev、EmuDev 等)对 LLM 的使用持强烈排斥态度。作者指出,这些社区的核心价值在于艰难地掌握领域知识的过程本身——运行代码只是附带成果,而 LLM 的使用被视为一种“作弊”,剥夺了真正的学习与 craftsmanship。尽管 LLM 在专家手中可以作为杠杆,但社区更看重的是通过长期活动、分享优雅代码、展示好奇心与深度知识来赢得尊重。早期尝试使用 LLM 的人往往缺乏深入理解,加上社区内的激烈守门行为,导致氛围迅速恶化。
https://news.ycombinator.com/item?id=49187061
https://research.meta.ai/blog/introducing-muse-code-and-muse-spark-1-2
Meta AI 发布了 Muse Code(测试版),这是一个由最新模型 Muse Spark 1.2 驱动的终端编码代理,标志着向更强大的前沿模型迈进。Muse Code 能处理大型代码库中的复杂软件工程任务,包括规划变更、编写代码和验证结果,并可协调多个持久运行的异步后台代理,减少延迟和人工干预。其运行时采用本地事件日志,确保回放精确且崩溃后可恢复,适合长时间运行的任务。内置技能包括 /plan(制定需批准的计划)、/grill(压力测试计划)和 /goal(推动目标完成)。
Muse Spark 1.2 是聚焦编码的模型更新,显著提升了代码生成、复杂调试、代码库理解和端到端开发工作流能力,同时保持通用代理优势。该模型与 Muse Code 联合训练,针对长时任务进行了大量训练,包括整库生成和自动研究,并利用规划、目标条件和上下文压缩维持进度。此外,通过自我改进循环,使用旧模型生成训练数据,提升了复杂指令遵循的准确性。
案例研究显示,Muse Spark 1.2 能在 GPU 内核优化中持续改进,通过代理环境编写、编译、剖析并优化内核性能。Muse Spark 1.2 现已通过 Muse Code 和 Meta 模型 API 向全球开放。
https://news.ycombinator.com/item?id=49187575
https://www.wired.com/story/meta-ran-ads-that-contained-ai-generated-child-sexual-abuse-imagery/
Meta 在 Facebook、Instagram 等平台投放了超过 50 条包含 AI 生成儿童性虐待图像的广告,部分广告本周仍在运行。文章指出这些广告内容涉及儿童性虐待描述,并引发了对 AI 技术缺乏道德准则的批评。评论者认为 Meta 和扎克伯格应为此负责,并呼吁对分发此类内容的行为进行法律追责。
https://news.ycombinator.com/item?id=49187977
纳什维尔市议会以 27 比 5 的投票结果,批准市长弗雷德·奥康奈尔的提案,授权市政府通过征用权收购动物园附近的一处地产,以阻止数据中心项目。该地块由开发商 DC Blox 持有,面积 23 英亩,计划建设一座 69,220 平方英尺的单层数据中心,投资超 7 亿美元。
动物园方面以噪音、光照、电力资源及环境影响为由强烈反对,并发起请愿,已收集超 50 万签名。反对者包括音乐明星布拉德·佩斯利、雪莉·克劳和杰克·怀特。市议会还通过了新的分区限制,并实施暂停新许可审批至 12 月 1 日。
目前全美超过 200 个社区和至少 14 个州正在考虑类似限制措施,纽约州也已暂停数据中心许可一年以研究其影响。DC Blox 表示暂无新进展可公布。
https://news.ycombinator.com/item?id=49191624
https://news.ycombinator.com/item?id=49193394
Programming has five phases effectively:
You figure out what problem to solve.
You figure out HOW to solve the problem.
You actually implement the solution.
You see the solution work, for yourself.
You ship/deploy/publish the program. This means you see people be happy users and/or you get paid for it and so on.
If you’re an entrepreneur type, you probably enjoy the first and last steps most, and you see steps 2-4 as mostly a chore. If you’re a tinkerer, you don’t care much for 1 and 5, and you see 2-4 as the whole point of programming. I’m a tinkerer. I’d be happy to just write code and throw it away. Coding is like solving sudokus. I could skip steps 1 and 5 forever. I don’t ever need to show any code to anyone. In fact, most of the time when programming I do steps 2 and 3 and even skip 4. I don’t even finish! I work weeks on something until I lose interest, and I know that in order to even run it, it would be several more weeks. A PoC is enough. Or just a half one. It’s just code-to-structure-thoughts, not to create anything finished.
The 5 phases look kind of symmetric. The outermost layer (1 and 5) are the entrepreneurial steps. If you’re a product owner or CEO, you might work strictly at steps 1,5. Then steps 2-4 are the managerial/architectural steps. If you’re a very senior IC at a large company, you might work at this level, without actually doing much coding. Only the inner most step (3) is the manual creation of source code. Even though it’s 5 different phases, it’s just “3 layers” of programming.
The problem as I see it is that I enjoy step 3. And LLMs are good at step 3 almost exclusively. So they just pick the best bit of this dish, and leave me with the rest.
If you’re an entrepreneurial type, the LLM appears to take the worst bit of the work from you. Great.
alkonaut
编程实际上有五个阶段:
你想清楚要解决什么问题。
你想清楚如何解决这个问题。
你真正实现这个解决方案。
你亲眼看到解决方案能跑起来。
你发布/部署/上线这个程序。这意味着你看到用户用得开心,和/或你因此赚到了钱,等等。
如果你是创业型的人,你大概最喜欢第一和最后一步,而把第2到第4步看作主要是苦差事。如果你是个喜欢捣鼓的人,你不太在乎第1和第5步,而是把第2到第4步当作编程的全部意义。我就是那种爱捣鼓的人。我只要能写代码然后扔掉就很开心。写代码就像解数独。我可以永远跳过第1步和第5步。我根本不需要向任何人展示代码。事实上,大多数时候我编程只做第2和第3步,甚至跳过第4步。我甚至不会完成!我会在某个东西上折腾几周直到失去兴趣,而且我知道,就算要把它跑起来,还得再花好几周。一个概念验证就够了,甚至半个都行。写代码只是为了把想法结构化,而不是为了做出什么成品。
这五个阶段看起来有点对称。最外层(第1和第5步)是创业性的步骤。如果你是产品负责人或CEO,你可能严格在第1和第5步工作。而第2到第4步是管理/架构层面的步骤。如果你是大公司里非常资深的个人贡献者,你可能就工作在这个层面,实际上不用写太多代码。只有最内层的第3步才是手动编写源代码。虽然说是五个不同阶段,但编程其实只有“三层”。
在我看来,问题在于我喜欢第3步。而大语言模型几乎只擅长第3步。所以它们只是把这道菜里最好吃的那一口挑走了,把剩下的留给我。
如果你是创业型的人,大语言模型看起来是从你手里接过了最糟糕的那部分活儿。很好。
https://news.ycombinator.com/item?id=49189872
This funny post reads surreal, but it may carry some truth: https://x.com/signulll/status/2067446889956430273?lang=en
hintymad
这个有趣的帖子读起来很超现实,但可能有些道理:https://x.com/signulll/status/2067446889956430273?lang=en
https://news.ycombinator.com/item?id=49185997
So, in last several months, all the prominent names Google lost: Demis Hassabis (technically still with google but these things are usually presented with a spin), Jeff Dean, Sanjay Ghemawat, Oriol Vinyals, Quoc Le, Noam Shazeer, John Jumper, Jonas Adler, Alexander Pritzel, David Silver, Denny Zhou, Fernando Pereira, Alex Turner
And all the prominent names Google gained: NULL
Combined with no gemini frontier GA release in about 14 months. You have to have created an environment pretty hostile to innovation for this to happen
GodelNumbering
在过去几个月里,谷歌流失的所有知名人士:戴密斯·哈萨比斯(严格来说仍在谷歌,但这类事情通常会被粉饰)、杰夫·迪恩、桑杰·格马沃特、奥里奥尔·维尼亚尔斯、郭乐、诺姆·沙泽尔、约翰·詹珀、乔纳斯·阿德勒、亚历山大·普里策尔、大卫·西尔弗、周登尼、费尔南多·佩雷拉、亚历克斯·特纳
以及谷歌获得的所有知名人士:无
再加上大约14个月内没有发布Gemini Frontier GA版本。你不得不承认,他们创造了一个对创新极为不利的环境。
https://news.ycombinator.com/item?id=49196345
It’s kinda funny there is still software coming out whose security model is “constantly ask the user for permission, and hope they never make a mistake”.
It’s been tried so many times before, and it never worked.
continuational
有点好笑的是,现在居然还有软件的安全模型是“不断向用户请求权限,并指望他们永远不会犯错”。这种模式以前试过很多次了,从来就没起过作用。
https://news.ycombinator.com/item?id=49191323
The web experience has improved? You can’t even view a thread any more without being logged in or using a 3rd party site. They’ve actively made the web experience unusable for most people.
But considering they’ve also drastically reduced the quality of content on the platform, one could say they’ve done casual users a favor.
solid_fuel
网页体验有所改善?你现在不登录或不使用第三方网站甚至无法查看任何帖子。他们主动让网页体验对大多数人来说变得不可用。
但考虑到他们还大幅降低了平台上的内容质量,可以说他们也算是帮了普通用户一个忙。
https://news.ycombinator.com/item?id=49188206
They make many bold promises, but their core goal is neatly encapsulated on the website:
“Imagine a future where a handful of people can conduct scientific research and engineering tasks much more rapidly, and with higher quality, than massive teams of scientists and engineers do today.”
This is the same goal as every other AI company out there. Automate away the human employees and let a small number of “people” (note that they do not say scientists or engineers for this part) take the credit and financial rewards for every good thing this human-free system produces.
beloch
他们做出了许多大胆的承诺,但其核心目标在网站上被简洁地概括为:
“想象一个未来,少数人能够以比当今庞大的科学家和工程师团队更快的速度、更高的质量来进行科学研究和工程任务。”
这和所有其他AI公司的目标如出一辙:把人类员工自动化掉,然后让一小撮“人”(注意,这部分他们说的可不是科学家或工程师)去占有这个无人工系统所产生的一切好处的功劳和经济回报。
https://news.ycombinator.com/item?id=49191814
Because people that enjoy programming for programming’s sake don’t want an LLM to do the programming for them - isn’t it obvious? It’s just like with any other hobby, people who like car racing created rules that force you to drive yourself, even though they’d get faster lap times with electronic driver aids. People who enjoy grappling created rules that force you to grapple, even though striking could win a fight faster. People that enjoy chess don’t allow you to bring a computer to the chess table. The list is endless and the people that enjoy programming will create rules that force you to program, why shouldn’t they? It’s their hobby.
Schnitz
因为那些纯粹为了编程而享受编程的人,并不想让大语言模型替他们编程——这不是明摆着的吗?这跟任何其他爱好一样,喜欢赛车的人制定了规则,强迫你必须自己驾驶,尽管用电子驾驶辅助能跑出更快的圈速。喜欢摔跤的人制定了规则,强迫你必须摔跤,尽管出拳能更快赢得打斗。喜欢下棋的人不允许你把电脑带到棋盘旁。这样的例子数不胜数,而喜欢编程的人也会制定规则,强迫你必须自己编程,他们为什么不该这样呢?这是他们的爱好。
https://news.ycombinator.com/item?id=49187716
Zed should focus on basics. When it focuses on basics Zed is good.
https://github.com/zed-industries/zed/discussions/54150 failing to show newly created files and declining to provide a refresh button, instead adding a polling backend, breaks Zed on WSL.
Why a new version control system? Why not git, jj, or another existing system?
NoDodgeQuestion
Zed应该专注于基础功能。当它专注于基础功能时,Zed就表现不错。
https://github.com/zed-industries/zed/discussions/54150 无法显示新创建的文件,又拒绝提供刷新按钮,反而添加了一个轮询后端,这在WSL上会让Zed崩溃。
为什么要开发一个新的版本控制系统?为什么不用git、jj或其他现有系统?
https://news.ycombinator.com/item?id=49185267
It seems like the real news is Jeff and Sanjay are leaving Google, and Demis is effectively replacing Jeff as Chief Scientist for all of Alphabet.
The bigger deal is the departure of Jeff and Sanjay, rather than Demis moving into a different role.
ra7
似乎真正的新闻是Jeff和Sanjay要离开谷歌,而Demis实际上取代Jeff成为整个Alphabet的首席科学家。
更重要的事情是Jeff和Sanjay的离开,而不是Demis换了个角色。
https://news.ycombinator.com/item?id=49192730
The reputation of datacenters precludes them due to some high-profile abuses of locals, as well as cities overcommitting brittle and unprepared infrastructure to them.
Consumer energy prices in 13+ states hiked because the grid was unable to meet the sudden demand. SpaceX’s datacenter in Memphis used gas turbines to rapidly deploy Colossus before the grid could be upgraded, which resulted in an increase in local air emissions that was technically illegal. There are two Meta datacenters that oneshotted their local water supplies during construction and tests, resulting in sediment and bacteria going through consumers’ taps. NV Energy cut off an entire town of residential properties because datacenters were more profitable, in a higher-stakes parallel to the ongoing RAM woes in consumer electronics.
There’s been more than enough well-justified association between “datacenter” and “abuses local resources” that I think nobody should really be surprised at the backlash toward them. The backlash is so furious that it’s forming a rare example of a bipartisan, populist consensus on any given political issue.
ashleyn
数据中心的名声因一些引人注目的滥用当地资源事件,以及城市将脆弱且准备不足的基础设施过度承诺给它们而受损。
超过13个州的居民电价上涨,因为电网无法满足突然增长的用电需求。SpaceX在孟菲斯的数据中心在电网升级前使用燃气轮机快速部署了Colossus,导致当地空气排放增加,这从技术上讲是违法的。两个Meta的数据中心在建设和测试期间曾一次性耗尽当地供水,导致沉积物和细菌进入居民的自来水。NV Energy公司切断了整个城镇的住宅供电,因为数据中心更有利可图,这与消费电子领域持续的RAM短缺问题形成了更高风险下的相似局面。
“数据中心”与“滥用当地资源”之间的关联已经足够充分且合理,我认为没有人应对此引发的反对声感到意外。这种反对情绪如此激烈,以至于形成了在任何政治议题上都罕见的跨党派、民粹主义的共识。
https://news.ycombinator.com/item?id=49191980
I don’t know what DC Blox was going to put in this particular datacenter, but they are a “traditional” colocation datacenter company. Maybe this one would’ve been superdense GPUs with a jet engine as electrical generator causing all kinds of noise, but it’s also just as likely this would’ve been a project that 5 years ago no one would’ve batted an eye over, and wouldn’t cause any issues in the local neighborhood
That’s the unfortunate thing about the current climate around datacenters. Everyone hates AI, and Elon absolutely screwed Memphis, and everyone thinks anything with more than 10 computers in it is going to be an xAI monstrosity. There’s still a need for growth in just traditional datacenters hosting things like your email and Netflix, not just AI, and those projects are going to get swallowed up in the AI hubbub
mixdup
我不知道DC Blox原本打算在这个数据中心里放什么,但他们是一家“传统”的托管数据中心公司。也许这个项目原本会配备超级密集的GPU,并用喷气发动机来发电,产生各种噪音,但同样有可能的是,这个项目在五年前根本不会引起任何人的注意,也不会给当地社区带来任何麻烦。
这就是当前数据中心环境令人遗憾的地方。每个人都讨厌AI,而埃隆绝对把孟菲斯搞砸了,现在所有人都认为只要里面有超过10台电脑的东西就会变成xAI的怪物。传统数据中心(比如托管你的电子邮件和Netflix的内容)仍然需要增长,而不仅仅是AI,但这些项目将会被AI的喧嚣所吞没。
https://news.ycombinator.com/item?id=49185385
OpenAI and Anthropic have the freedom to do absolutely insane things like negligently hack other companies. It would be stock price suicide if anything even remotely happened with Google.
xnx
OpenAI和Anthropic可以自由地做出极其疯狂的事情,比如因疏忽而入侵其他公司。如果谷歌发生任何类似的事情,那将是股价的自杀。
https://news.ycombinator.com/item?id=49194543
My favorite Youtube channel.
Please watch the latest https://www.youtube.com/watch?v=HDVcfSYcxc0. It’s an interview with an old botanist, who shares his wisdom and philosophy, remains engaged after decades of witnessing first hand America’s nature getting destroyed. 2 main points: - destroyed old-growth forests, grassland, remnant prairies takes millennia to be back to what they were, not decades. Insects and other pollinators don’t just reappear - water availability for plants comes in many forms, all of which disappear with tillage.
Joey is an acquired taste, but if you look past the strong and odd personality (the persona?), his channel is a gold mine. But this time he was (rightfully) impressed by Gerry (Gerould) Wilhelm’s wealth of knowledge.
jojocool0501
我最喜欢的YouTube频道。
请观看最新的视频https://www.youtube.com/watch?v=HDVcfSYcxc0。这是一位老植物学家的访谈,他分享了自己的智慧与哲学,在亲眼目睹美国自然被破坏数十年后,仍然保持着参与。两个要点:- 被破坏的原始森林、草原、残留草原需要数千年才能恢复原状,而非数十年。昆虫和其他传粉者不会凭空重现 - 植物所需的水分有多种形式,所有这些都会随着耕作而消失。
Joey是一个需要慢慢品味的人,但如果你忽略他强烈而古怪的个性(或人设?),他的频道就是个宝藏。但这次他(理所应当地)被Gerry (Gerould) Wilhelm的渊博知识所震撼。
https://news.ycombinator.com/item?id=49185577
From Jeff’s twitter post:
Our general approach is to automate the experimental loop. We think this approach is broadly applicable across many different fields of science and engineering. We’ll initially focus on ML research and engineering, but believe the approach can help with important subproblems in nearly every one of the fourteen <at>NAE Grand Challenge problems. We think doing this well requires strong expertise in machine learning as well as large-scale systems.
See also: https://www.nae.edu/20782/grand-challenges-project
Those 14 are:
NAE Grand Challenges for Engineering
Make Solar Energy Economical
Provide Energy from Fusion
Develop Carbon Sequestration Methods
Manage the Nitrogen Cycle
Provide Access to Clean Water
Restore and Improve Urban Infrastructure
Advance Health Informatics
Engineer Better Medicines
Reverse Engineer the Brain
Prevent Nuclear Terror
Secure Cyberspace
Enhance Virtual Reality
Advance Personalized Learning
Engineer the Tools of Scientific Discovery
cjbarber
来自杰夫的推特帖子:
我们的总体方法是自动化实验循环。我们认为这种方法广泛适用于科学和工程的许多不同领域。我们将首先专注于机器学习的研究与工程,但相信这种方法能帮助解决美国国家工程院(NAE)十四大挑战项目中几乎每一个领域的重要子问题。我们认为要做好这一点,需要具备机器学习和大规模系统的深厚专长。
另见:https://www.nae.edu/20782/grand-challenges-project
这十四项是:
NAE工程领域的重大挑战
https://news.ycombinator.com/item?id=49188724
Sanjay (who just joined Twitter)
Clarification: this comment is saying Sanjay Ghemawat joined Twitter as a user recently (new account @Sanjay_Ghemawat as of July 2026), as opposed to Sanjay working for Twitter the company.
omoikane
Sanjay(刚加入推特)
澄清:这条评论是说Sanjay Ghemawat最近注册了推特账号(截至2026年7月的新账号@Sanjay_Ghemawat),而不是说Sanjay为推特工作。
https://news.ycombinator.com/item?id=49190589
https://xcancel.com/signulll/status/2067446889956430273
JumpCrisscross
这是典型的心理投射,更别说还带有那种卑微的自我贬低。连一个完整的句子都拼不出来,难怪得靠陌生人的食物才能活下去。
https://news.ycombinator.com/item?id=49187427
A hobby is something you enjoy the process of doing, not just the end result. Everybody likes a clean home, but cleaning is rarely a hobby. LLMs expedite achieving the end result. Take from that what you will.
QuantumNoodle
爱好是你享受过程本身的事情,而不仅仅是最终结果。每个人都喜欢干净的家,但打扫卫生很少成为爱好。大语言模型只是加速了最终结果的实现。你从中能悟到什么,就各凭心意了。
2026-08-06 07:37:13
- 斯蒂芬·沃尔夫拉姆悼念其数学家妻子埃莉斯·考利,回忆36年共同生活与思想碰撞,深感悲痛。
- “发现循环”项目旨在通过AI自动化科学实验循环,由Jeff Dean等专家领导,加速科研进程。
- Pi编码工具以极简设计著称,在真实任务中表现优异,上下文管理紧凑,可扩展性强。
- 新墨西哥州医疗运输机坠毁可能与军事GPS干扰有关,凸显无人机时代民用航空安全风险。
- Cloudflare OS是一个开放平台,整合全球网络能力,简化AI智能体、应用和后台任务的开发部署。
- Google DeepMind重大人事调整:Demis Hassabis转任主席,Jeff Dean离职创办新公司。
- 明尼苏达州威诺纳市警局全部8个Flock车牌识别摄像头被锯断盗走,总损失约24000美元。
- Demis Hassabis卸任Google DeepMind CEO转任董事长,多名AI同事离职创办新公司。
- 作者收到看似钓鱼的FedEx短信,经官方渠道核实实为合法通知,提醒应通过官方渠道验证。
- Waymo全自动驾驶出租车服务在达拉斯向所有用户开放,计划扩展至机场和高速公路。
这是一篇悼念文章,作者斯蒂芬·沃尔夫拉姆(Stephen Wolfram)为纪念其妻子埃莉斯·考利(Elise Cawley,1961–2026)而作。
文章讲述了埃莉斯因突发心血管事件去世,两人共同生活了 36 年。埃莉斯是一位才华横溢的纯数学家,性格聪慧、热情,追求真理与美。她不仅在数学领域有深厚造诣,还热爱室内设计、建筑和美学,并致力于为家庭创造温暖而富有设计感的家。
作者回忆了两人相识于 1990 年,以及 36 年间日常的交流与思想碰撞。埃莉斯对数学基础有独到见解,认为数学本质上是人文的,而非纯粹的形式化。她曾指出数学的基本抽象是“点”与“数”,而物理与数学的主要区别在于引入了“时间”概念。
文章还提到,埃莉斯既有强烈的自信,又保持谦逊,善于洞察事物的本质。她与作者在思维方式上互补,共同珍视主题性思考与美学。她的离世让作者深感悲痛,并希望通过此文让世人了解这位非凡的女性。
https://news.ycombinator.com/item?id=49173165
https://www.discoveryloop.com/
这是一个关于 Discovery Loop 项目的介绍页面。
该项目旨在通过自动化科学实验循环,解决当前科学发现速度缓慢的问题。其核心是利用前沿 AI 模型和大规模计算基础设施,自动完成实验的提出、执行和结果评估,从而并行运行数千个实验,大幅缩短迭代时间。
核心方法:
团队背景: 创始团队包括 Jeff Dean、Sanjay Ghemawat、Quoc Le 和 Oriol Vinyals。他们拥有数十年的深度合作经验,共同开创了大规模计算,并领导创建了 Google 搜索、MapReduce、TensorFlow、TPU、AlphaFold、Gemini 等关键基础设施与 AI 突破。
未来展望: 项目希望打造一个精干的小团队,实现“少数人能以更高效率和质量完成今天庞大团队才能完成的科研任务”的未来。
https://news.ycombinator.com/item?id=49184960
https://earendil.com/posts/pi-autoresearch-and-databricks/
Pi 是一款极简且高性能的编码工具,核心设计理念是“有意的极简主义”。它开箱仅提供 4 个工具,系统提示和工具定义总计不到 1000 个 token,旨在用最基础的功能完成大部分工作,用户可按需扩展。
Databricks 的研究表明,Pi 在真实编码任务中表现优异。当与 Opus 4.8 模型配合时,Pi 的通过率最高,且成本显著低于 Claude Code 和 Codex。Pi 的优势在于“上下文纪律”——每轮发送的上下文量约为其他工具的 1/3,管理更紧凑,任务完成所需轮次更少。
Shopify 基于 Pi 构建了“Autoresearch”扩展,这是一个用于编码代理优化的自主循环。它能自动运行实验来优化代码,例如使单元测试运行速度提升 300 倍,React 组件挂载速度提升 20%。这体现了 Pi 的可扩展性:不预装所有功能,而是让用户轻松构建自己的工具。
Pi 的极简设计在当前环境下更具优势。前沿模型已能很好地理解终端环境,因此工具的关键不再是“原生性”,而是如何高效管理上下文、避免冗余。Pi 通过更少的提示开销、更低的运行成本和更少的抽象层,提供了更清洁的模型接口。对于本地模型,Pi 的上下文纪律尤其有价值,能避免长时间重新预填充。
https://news.ycombinator.com/item?id=49176038
https://www.wired.com/story/a-civilian-plane-crashed-in-new-mexico-was-the-militarys-tech-to-blame/
今年 5 月,一架双引擎“比奇空中国王”医疗运输机从新墨西哥州罗斯威尔起飞,前往鲁伊多索接病人。天气晴朗,航程约 60 英里。然而,这架搭载两名飞行员和两名护士的飞机最终坠毁。
文章探讨了这起事故的潜在原因:军事无人机训练导致的 GPS 干扰(代号 NAVFEST)是否难辞其咎。美国联邦航空管理局曾警告,5 月 12 日至 18 日期间 GPS 可能受到影响,而事故发生在 5 月 14 日。
有飞行员评论指出,事故主要责任在于飞行员过度依赖 GPS,忽视了基本的飞行技能和地形图检查。军方虽发出了相关通知,但可能存在信息混淆。这起事故凸显了无人机战时代,民用航空面临的新的安全风险。
https://news.ycombinator.com/item?id=49181099
https://blog.cloudflare.com/cloudflare-os/
好的,这是根据您提供的 HTML 内容生成的网页摘要:
该网页是 Cloudflare 的官方博客文章,标题为“Cloudflare One 数据保护套件”。文章主要介绍了 Cloudflare 推出的一套新的数据保护解决方案,旨在帮助企业在使用 AI 工具和 SaaS 应用时,防止敏感数据泄露。
核心内容包括:
https://news.ycombinator.com/item?id=49182996
https://blog.google/company-news/inside-google/message-ceo/next-chapter-ai-momentum/
Google 和 Alphabet 的 CEO 桑达尔·皮查伊宣布了 Google DeepMind 团队的重大人事调整,旨在加速 AI 发展并聚焦通用人工智能(AGI)的未来。
核心变动包括:DeepMind 联合创始人德米斯·哈萨比斯将担任 Google DeepMind 主席兼 Alphabet 首席科学家,专注于 AGI 和科学领域的战略工作。现任首席技术官科拉伊·卡武克丘奥卢将升任 Google DeepMind 高级副总裁,负责 Gemini 模型开发、前沿 AI 研究及应用团队。此外,资深高管杰夫·迪恩将与桑杰·格玛瓦特共同创立一家公益公司,专注于机器学习、科学和工程领域的加速发现。
https://news.ycombinator.com/item?id=49184755
明尼苏达州威诺纳市警局(Winona Police Department)部署的全部 8 个 Flock 车牌识别摄像头于 8 月 1 日被人锯断盗走。巡逻警官发现摄像头 24 小时未发送警报后进行检查,发现摄像头从杆上被割断,仅留下杆子。此外,布法罗县在密西西比河大桥上的 2 个 Flock 摄像头也以同样方式被盗。
每个摄像头价值约 3000 美元,威诺纳警局直接损失约 24000 美元。这些摄像头位于城市主要高速公路出入口,可自动捕捉车牌及车辆的品牌、型号、颜色,警方用于调查肇事逃逸和寻找失踪人员。摄像头数据由警局掌握,30 天后永久删除。
此类盗窃已成为全国趋势,多个社区的 Flock 摄像头曾遭破坏。该技术也因扩大政府监控及数据共享问题受到民权倡导者批评。目前案件仍在调查中,尚无嫌疑人,警方呼吁知情者提供线索。
https://news.ycombinator.com/item?id=49171656
https://www.axios.com/2026/08/05/google-deepmind-demis-hassabis-ai
Google DeepMind CEO Demis Hassabis 将卸任,转任该部门董事长。首席科学家 Jeff Dean 及多位 Google AI 同事离职创办新公司 Discovery Loop,Google 将对其进行投资。Hassabis 还将担任 Alphabet 首席科学家,并继续领导 AI 药物发现公司 Isomorphic Labs。Koray Kavukcuoglu 将接任 DeepMind 高级副总裁。
此次人事变动正值 Google AI 组织面临深刻变革,在追赶 OpenAI 和 Anthropic 的过程中遇到挑战。Google 股价下跌超 4%,多名顶尖研究人员(包括 Gemini 联合负责人)已离职。Hassabis 表示将更专注于 AGI 的科学方向,而 Jeff Dean 则希望在新公司中更自由地探索 AI 科学发现。
https://news.ycombinator.com/item?id=49184757
https://www.troyhunt.com/thanks-fedex-this-is-why-we-keep-getting-phished/
作者 Troy Hunt 收到一条疑似钓鱼的 FedEx 短信,要求支付关税和税款。短信中有多个可疑特征:拼写错误(FedEx 写成 FedEx-Exp)、短而不规范的追踪号、紧急语气、大写字母异常、未标明货币、链接指向 bpoint.com.au 而非 FedEx 官网等。87% 的网友认为这是诈骗。但作者确实在等待一个从海外寄来的 3D 打印机,可能涉及关税。他尝试通过 FedEx 官网验证,但官网未找到相关税务信息。随后他拨打 FedEx 客服电话,陷入自动语音循环,最终接通人工客服,确认包裹需缴税。三天后,他收到 FedEx 官方邮件,附有完整发票,证明短信真实。文章指出,尽管短信看似可疑,但实际是合法通知,提醒人们不要仅凭表面判断,应通过官方渠道核实。
https://news.ycombinator.com/item?id=49175192
https://waymo.com/blog/shorts/dallas-open-to-all/
Waymo 宣布,自 2026 年 8 月 4 日起,达拉斯所有用户均可通过下载 Waymo 应用体验全自动驾驶出租车服务。自 2 月开放以来,已有近 15 万用户从兴趣列表加入,如今服务面向所有人开放。
Waymo 正在达拉斯爱田机场航站楼进行全自动驾驶测试,并计划很快在达拉斯高速公路上开展测试,完成后将向公众开放这些路线。
癫痫基金会德州分会首席执行官表示,Waymo 自动驾驶汽车为因医疗条件无法驾驶的人群提供了安全、独立的出行方式,是变革性的进步。
https://news.ycombinator.com/item?id=49172836
https://news.ycombinator.com/item?id=49176894
The net result of what this site is doing seems to result in one of two effects when I open the sidebar, randomly:
the sidebar covers the content, while the content has blank space to the right, or
the content moves off the right edge of the screen, leaving a large blank space to its left.
Both of these are wrong. If I have a sidebar open, the site is now narrower , stop trying to be clever.
My immediate reaction to this is why is the browser giving the site this information, and could we stop.
JoshTriplett
每次打开侧边栏时,这个网站的操作最终似乎会导致以下两种情况之一随机出现:
侧边栏覆盖了内容,而内容右侧出现空白区域,或者
内容移出屏幕右边缘,在其左侧留下大片空白。
这两种情况都是错误的。如果我打开了侧边栏,网站应该变窄,别自作聪明。
我的第一反应是:为什么浏览器要给网站提供这些信息,以及我们能否阻止这种行为。
https://news.ycombinator.com/item?id=49176601
This might be some kind of a weird requirement for a piece of art, but I would literally never expect a website to center a div according to the browser window instead of the viewport. It just looks wrong and feels wrong.
kccqzy
这可能是对一件艺术作品的一种奇怪要求,但我真的从未想过一个网站会按照浏览器窗口而不是视口来居中一个div。这看起来不对,感觉也不对。
https://news.ycombinator.com/item?id=49174158
An amazingly detailed tribute, you’d have to think this was taken from a journal, in which case it’s impressive to keep such a record of their life, or maybe Stephen just has a fantastic memory.
Even through that detail you can really feel how much he loved and admired her.
Somehow in the shock of what has happened it feels as if I just met Elise, and now she is gone.
A reminder that time’s cruelty is making the wonderful feel shorter than the painful.
My only solace in this tragedy is that the end came instantly, after a day filled with nothing but joy.
Hopefully we’re all lucky enough to be allowed to rest without suffering while surrounded by loved ones.
cube00
一份惊人详尽的悼念,你会觉得这像是从日记里摘录出来的——若果真如此,能如此细致记录他人生活已令人赞叹;又或者,斯蒂芬只是拥有超凡的记忆力。
即便在那些细节中,你也能真切感受到他有多么爱她、敬慕她。
在事发后的震惊中,恍然觉得我仿佛才刚遇见艾莉丝,而如今她已离去。
这提醒我们,时间的残酷在于,它让美好的时光比痛苦的更显短暂。
这场悲剧中我唯一的慰藉,是她在充满欢愉的一天后,瞬间安详离去。
但愿我们都能足够幸运,在亲人的环绕中无痛地安息。
https://news.ycombinator.com/item?id=49184701
This is the big problem with mass surveillance as a whole. With enough time, they’ll find some pretext to bring you in, even if the original crime they want you for is weak.
This will ruin lives, even if they’re eventually let go. Jobs lost, evicted from homes, cars towed, pets surrendered and put down.
Lives will be completely derailed, police and judges will not be held accountable. The wrongful arrestees will still be in debt to the local jail and court for various fees they racked up. Trust in the justice system will continue to degrade and more and more violence will be used against people to maintain order.
We hold police to high standards for search/arrest because they have the monopoly on violence. Even if they’re wrong, you can’t shoot back. Even if they’re slamming your loved ones into the pavement, breaking their teeth, you cannot kick the cop in the head or else they have permission to murder you both.
Mass surveillance is a loophole to this contract.
MSFT_Edging
大规模监控的整体问题就在于此。只要有足够的时间,他们总能找到某种借口把你抓进去,哪怕最初想定你的罪名根本站不住脚。
即便最终被释放,生活也会因此被毁:丢掉工作、被赶出住房、汽车被拖走、宠物被遗弃甚至安乐死。
人生将彻底脱轨,而警察和法官却不会承担任何责任。被错误逮捕的人还得向当地监狱和法院支付他们积欠的各种费用。公众对司法体系的信任会持续恶化,为了维持秩序,针对民众的暴力手段会越来越多。
我们对警察的搜查与逮捕设置高标准,是因为他们垄断了暴力权力。即便他们是错的,你也不能开枪还击。即便他们把你的亲人猛摔在人行道上、打碎他们的牙齿,你也不能一脚踢向警察的头,否则他们就有权把你俩都杀了。
大规模监控正是这一社会契约的漏洞所在。
https://news.ycombinator.com/item?id=49183266
I liked Kenton’s take on this: https://x.com/KentonVarda/status/2084990137180590572?s=20
Text from tweet:
Today we are releasing Cloudflare OS, a chatbot with connectors, just like every other tech company is doing.
Except actually, it’s different. This is a remake of Sandstorm[.]io, my startup from 10 years ago, except this time built on Cloudflare Workers (the platform I’ve spent the last 9 years building) and deeply leveraging AI. This is more or less the culmination of my secret 10-year master plan.
This is a full-on personal app vibe coding platform, in which the sandbox is so secure that you can pretty much go wild – the AI cannot introduce a significant security bug. We believe a company’s security team can feel comfortable giving non-technical users permission to vibe code and then sleep soundly at night.
How is that possible? It’s the Sandstorm security model, revisited. A “Gadget” is the same thing as a Sandstorm “Grain”: a fine-grained app instance. For example, if you have a document editor app, each document runs as a separate instance of the app, in a separate sandbox (one “Gadget”).
This means two things, both of which I think are Big Deals: 1. The platform can manage all access control, by controlling who can access the Gadget at all. There is no way the Gadget can accidentally leak itself to an attacker – even an attacker who has access to other Gadgets based on the same app. 2. Since everyone is running their own copy of the code, everyone can freely modify their copy of the code.
Think about #2 a bit more.
What if, when you wanted a new feature in the software you are using, you could just prompt your agent to add it?
This doesn’t work in the cloud Software-as-a-Service model, because you are not running your own copy of the app.
Sandstorm tried to change that 10 years ago, but the world wasn’t ready, because not enough people had the skills or patience to actually modify their software. AI has changed that. Now you just ask the agent – the same agent that you are using to help you interact with the Gadget can also modify the code of the Gadget.
And it is so fun.
rozenmd
我喜欢Kenton对此的看法:https://x.com/KentonVarda/status/2084990137180590572?s=20
推文内容:
今天我们发布了Cloudflare OS,一个带连接器的聊天机器人,就像其他所有科技公司正在做的一样。
但实际上,它有所不同。这是对Sandstorm[.]io的重制版,我十年前创办的公司,只不过这次构建在Cloudflare Workers(我过去九年一直构建的平台)之上,并深度利用人工智能。这或多或少是我秘密的十年大计划的巅峰之作。
这是一个完整的个人应用氛围编码平台,其中的沙箱非常安全,你几乎可以尽情发挥——人工智能不会引入重大安全漏洞。我们相信,一家公司的安全团队可以放心地允许非技术用户进行氛围编码,然后安然入睡。
这怎么可能?这是对Sandstorm安全模型的重新审视。一个"小工具"就像Sandstorm中的"颗粒":一个细粒度的应用实例。例如,如果你有一个文档编辑器应用,每个文档都作为该应用的一个独立实例运行,在独立的沙箱中(一个"小工具")。
这意味着两件事,我认为都很重要:1. 平台可以通过控制谁能访问该小工具来管理所有访问控制。小工具不可能意外将自己泄露给攻击者——即使攻击者可以访问基于同一应用的其他小工具。2. 由于每个人都在运行自己的代码副本,每个人都可以自由修改自己的代码副本。
再想想第2点。
如果你在使用软件时想要一个新功能,只需提示你的智能体添加它,会怎么样?
这在云的"软件即服务"模型中行不通,因为你没有运行自己的应用副本。
Sandstorm十年前试图改变这一点,但世界还没有准备好,因为没有足够的人有技能或耐心去实际修改他们的软件。人工智能改变了这一点。现在你只需告诉智能体——那个正在帮助你与小工具互动的智能体也可以修改小工具的代码。
而且这非常有趣。
https://news.ycombinator.com/item?id=49178337
With respect to you as a person and your ideas presented, this is one of the most car-brained posts I’ve ever read.
it’s way more efficient than building $100 billion in transit that will take 200 years to build.
What is efficient about a $13 highway widening project in Houston or the $24 billion (inflation-adjusted) and 15 years it cost to bury a highway in Boston? (It’s not even much of a new highway, it’s just underground now!)
I don’t think the average person realizes how little typical non-New York states spend on transit. Ohio’s ratio on highways:public transit spending is approximately 8:1, and that includes federal public transit money. Oh yeah, and Ohio outspends all its bordering neighbors per-capita.
I am here to tell you about this cutting-edge science-fiction technology that is cheap, efficient, highly flexible, and requires almost zero infrastructure investment: https://en.wikipedia.org/wiki/Bus
As a bonus, instead of funneling public money out of state to Waymo who then dumps into cars manufactured by Zeekr in China, the city bus actually employs local people!
It’s so fascinating that Arlington, Texas, a city with over 300,000 people in it, hasn’t yet been able to master the deep intricacies and mystical technology of the bus.
I’ve also noticed that you are talking about the impacts of transit and vehicles on real estate prices, e.g., a lack or excess of parking spaces. I think you can do something similar the other way around and investigate the real estate impact of permanent rapid transit stations being built. That paradigm is how Japan funds its entire public transit system as a profitable enterprise: allow the rail company to own and develop real estate in close proximity to the stations.
Grombobulous
尊重你作为个人以及你所提出的观点,但这篇帖子是我读过的最“车本位”的文章之一。
“这比花1000亿美元建个200年才能建成的公交系统高效多了。”
休斯顿那个130亿美元的高速公路拓宽项目,或者波士顿花了240亿美元(按通胀调整)、耗时15年才把一条高速埋入地下(其实根本不算什么新高速,只是把它转到地下了而已),到底哪里高效了?
我觉得一般人根本不知道普通非纽约州在公共交通上花的钱有多可怜。俄亥俄州在高速公路与公共交通上的支出比例大约是8:1,这还已经算上了联邦给公交的钱。对了,俄亥俄州的人均交通支出还超过了它所有接壤的邻州。
我来给你介绍一种前沿科幻技术,它便宜、高效、极其灵活,而且几乎不需要任何基础设施投资:https://en.wikipedia.org/wiki/Bus
额外的好处是,与其把公共资金送到州外给Waymo,然后Waymo再把钱投进中国极氪制造的汽车里,城际公交实际上雇佣的是本地人!
真有意思,得克萨斯州阿灵顿市——一个超过30万人口的城市——至今还没能掌握公交这种深邃精妙的神秘技术。
我还注意到你在谈论交通和车辆对房价的影响,比如停车位不足或过剩。其实你也可以反过来做类似的事,研究一下永久性快速公交站建成后对房地产的影响。这种模式正是日本把整个公共交通系统做成盈利企业的方式:允许铁路公司拥有并开发车站附近的房地产。
https://news.ycombinator.com/item?id=49174237
Surprisingly low hype for a company running perhaps the most advanced consumer-interactive robots. I’m quite fond of these machines and will periodically rescue them when someone has misused them (door not closed, most commonly). Hope they do well. Very good road participants.
arjie
对于一个运营着可能是最先进的消费级交互机器人的公司来说,关注度出奇地低。我非常喜欢这些机器,每当有人误用它们时(最常见的是没关好门),我都会定期去“解救”它们。希望他们能发展得好。非常出色的道路参与者。
https://news.ycombinator.com/item?id=49181565
When I used to fly in Florida in the 2005-2015ish timeframe, NOTAMs for GPS interference (it didn’t say that exactly but that’s what it meant) were an everyday occurrence along my path, it became boilerplate. I experienced it a few times. One moment the GPS thinks you’re plodding along next to Ocala, next moment it thinks I’ve teleported a thousand miles away, and then it’d just quit.
But that’s what we train for. That’s why aircraft and pilots have redundancy to work around any single possible point of failure and why you practice it until it becomes natural. Because of the NOTAM I’d already have a VOR tuned in, and once the GPS went inop I just fine tuned it to get a radial fix on the VOR. Also tried to never lose my bearings with my MK I Eyeballs either. No big deal.
I feel bad for everyone that lost their lives here, but this is why competence is critical. There’s a litany of errors the pilots made. And the article tries to pin a bit on the controller, but no, ultimate responsibility is on the pilot in command, FAR/AIM is clear on that.
If these allegedly instrument rated pilots had been properly drilled by their CFI’s and properly grilled during reviews and such this should’ve been an easy flight.
Also, “just ask Kennedy”. Wired bringing that into the conversation is perhaps relevant but not in the way I think they intended. JFK Jr lacked basic airmanship skills. His aircraft was 100% functional until the moment it impacted the ocean.
Some professions simply can’t be go-along-to-get-along social clubs. Safety culture centered on competence and performance is key.
mrngld
2005到2015年左右我在佛罗里达飞行时,航路上关于GPS干扰的航行通告(措辞不完全如此,但实际就是那个意思)几乎天天出现,都成了套话。我亲身经历过几次。前一秒GPS还认为你在奥卡拉附近慢慢飞行,下一秒就觉得你瞬移到了千里之外,接着就直接罢工了。
但这也是我们训练的意义所在。这就是为什么飞机和飞行员要有冗余设计来应对任何一个可能的故障点,以及为什么你要反复练习直到形成本能。因为看到了通告,我早就调好了VOR台,一旦GPS失效,我微调一下就能通过VOR获得径向线定位。同时我也尽量用肉眼保持方向感。没什么大不了的。
我为在这里遇难的所有人感到难过,但这也正是能力至关重要的原因。飞行员犯了一连串错误。文章试图让管制员承担部分责任,但不对,最终责任在于机长,联邦航空条例对此有明确界定。
如果这些声称有仪表等级的飞行员曾被他们的飞行教官严格训练,并在复训中被严格考核,这次飞行本应轻松完成。
另外,“问问肯尼迪就知道了”——把这句话扯进讨论或许有点关联,但恐怕不是《连线》杂志想表达的意思。小肯尼迪缺乏基本的飞行技巧。他的飞机直到坠海那一刻都完全正常运行。
有些职业不能成为"随大流混日子"的社交俱乐部。以能力和表现为核心的安全文化才是关键。
https://news.ycombinator.com/item?id=49180965
Would be interesting to join but have to get bit more karma :)
toivo
加入会很有趣,但得先攒点 karma :)
https://news.ycombinator.com/item?id=49188174
The AI is no longer “other”
My AI is a natural extension of me. It feels like a sixth sense and another limb. I wonder about a problem, feel as though I’m literally surfing the web, see glimpses of the websites, get flashes of intuition about the problem, and ultimately derive the answer.
We call this AI psychosis.
thatmf
人工智能不再是“他者”
我的人工智能是我自然的延伸。它就像是第六感和另一只肢体。我对某个问题感到好奇,仿佛真的在网络上冲浪,瞥见网页的片段,对问题产生直觉的闪念,并最终推导出答案。
我们称之为AI精神病态。
https://news.ycombinator.com/item?id=49177020
God, I love how Americans will do anything but build public transport
m_a_g
老天,我真服了美国人什么都愿意做,就是不肯建公共交通。
https://news.ycombinator.com/item?id=49176828
The longer background is that Christian Ude our mayor until 2014 furthered the LiMux[1] project - migrating more than 14000 PCs in public administration to Linux.
He got personal visits from both Ballmer and Gates that pressured him to stay with Microsoft[2] but did not balk.
Unfortunately his successor caved to Microsoft’s siren song abolished LiMux and Microsoft got a nice campus in the city (they used to be at the outskirts).
Since May this year we have the young and energetic Dominik Krause and there is hope for change again.
[1] https://en.wikipedia.org/wiki/LiMux
[2] https://www.linux-magazin.de/ausgaben/2019/10/interview-2/
weinzierl
更长的背景是,我们的市长克里斯蒂安·乌德(任期至2014年)推动了LiMux[1]项目——将公共行政部门超过14000台个人电脑迁移到Linux系统。
他曾受到鲍尔默和盖茨本人的亲自拜访,施压他继续使用微软[2],但他并未退缩。
不幸的是,他的继任者屈服于微软的诱惑,废除了LiMux项目,而微软则在市内获得了一处漂亮的园区(他们原本位于市郊)。
自今年五月起,我们有了年轻且充满活力的多米尼克·克劳斯,希望再次出现转机。
[1]
https://en.wikipedia.org/wiki/LiMux
[2]
https://www.linux-magazin.de/ausgaben/2019/10/interview-2/
https://news.ycombinator.com/item?id=49177686
This is a nonsensical goal.
If I drag a browser window to the right, making it cover the right-hand half of my monitor, it serves no useful purpose for “centered” content to migrate to the left edge of the window (the physical midpoint of the screen).
Centered means centered within the virtual viewport, not the physical screen.
Borealid
这是一个毫无意义的目标。
如果我将浏览器窗口拖到右侧,使其覆盖显示器右半部分,那么“居中”的内容却跑到窗口左边缘(屏幕的物理中点),这毫无用处。
居中是指相对于虚拟视口居中,而不是物理屏幕。
https://news.ycombinator.com/item?id=49186101
I run GPSJAM.org, have been studying and tracking the effects of GPS interference on aviation for the past 5 years, and I was a source for this article. I’m not a pilot and I don’t work for any aviation agency.
I’m very curious to see what the final NTSB report says, but their preliminary report along with commentary from other analysts seems to show that this was a crew that made bad choices and died because of it. It also seems to show that GPS interference was a contributing factor to the accident, and that they’d likely be alive if the U.S. military hadn’t been jamming GPS.
I think pilots need to be prepared to fly without GNSS, and also it might be a bad idea for the military to regularly deny GPS to thousands of civilian aircraft with tens or hundreds of thousands of passengers.
Lack of GPS isn’t a critical safety issue, but it is a safety issue. GPS interference removes options, hurts situational awareness, disables or degrades other safety equipment (like TAWS, the Terrain Awareness and Warning System that is specifically designed to warn pilots that they’re about to hit a mountain), and adds to crew and ATC workload and distraction.
Saying that this situation is what pilots train for and they should have just used VORs and ILS reveals an overly macho, unsophisticated understanding of aviation safety. We’ve added GPS, ADS-B, TCAS, GPWS, TAWS, etc. because they make flight safer. The airline industry and governments are concerned about GPS interference because of the aviation safety issues–if you’re not concerned, you’re thinking about the problem in a very narrow way (" I would never have trouble if my GPS died.").
In fact they think it’s an urgent problem.
2024: “Aviation sector seeks urgent solutions for GPS interference” https://www.reuters.com/business/aerospace-defense/aviation-sector-seeks-urgent-solutions-gps-interference-2024-01-24/
From a 2023 presentation by Eurocontrol at the UN International Committee on GNSS/Interference Detection and Mitigation https://rntfnd.org/wp-content/uploads/Aviation-GNSS-interference-UN-ICG-WGS-IDM-ECTL-GNSS-RFI-SEP23.pdf :
• GPS Problem reports dominate over all other type of safety reports • Redundant systems are the only reason why aviation has been able to maintain normal operations despite GNSS RFI! • GNSS integrated into many systems; exact RFI impact difficult to predict, manufacturers had to issue aircraft specific guidance → Complexity and workload increase
Aviation Safety Impact • Aviation Safety is built on two main principles: • Trust your instruments • Follow standard operating procedure • GNSS RFI causes pilots to have to question both principles! • Chief Operations Officer of one major airline: Navigation is not my problem. My problem is “normalization of deviance”! • Incidents have occurred simply due to pilot distraction because of having to deal with too many system alerts The U.S. military jamming GPS of civilian aircraft is a situation where there have been several close calls before this incident and experts have been saying for years that if the military keeps doing this, even with NOTAMs, people will get hurt and then that prediction tragically came true.
From 2021, “FAA Files Reveal a Surprising Threat to Airline Safety: The U.S. Military GPS Tests” https://spectrum.ieee.org/faa-files-reveal-a-surprising-threat-to-airline-safety-the-us-militarys-gps-tests :
Early one morning last May, a commercial airliner was approaching El Paso International Airport, in West Texas, when a warning popped up in the cockpit: “GPS Position Lost.” The pilot contacted the airline’s operations center and received a report that the U.S. Army’s White Sands Missile Range, in South Central New Mexico, was disrupting the GPS signal. “We knew then that it was not an aircraft GPS fault,” the pilot wrote later.
The pilot missed an approach on one runway due to high winds, then came around to try again. “We were forced to Runway 04 with a predawn landing with no access to [an instrument landing] with vertical guidance,” the pilot wrote. “Runway 04…has a high CFIT threat due to the climbing terrain in the local area.” Also
This is far from the most worrying ASRS report involving GPS jamming. In August 2018, a passenger aircraft in Idaho, flying in smoky conditions, reportedly suffered GPS interference from military tests and was saved from crashing into a mountain only by the last-minute intervention of an air traffic controller. “Loss of life can happen because air traffic control and a flight crew believe their equipment are working as intended, but are in fact leading them into the side of the mountain,” wrote the controller. “Had [we] not noticed, that flight crew and the passengers would be dead. I have no doubt.” Here’s what an air traffic controller with 18 years of experience wrote about WSMR GPS interference in 2024: “Someone will get hurt ignoring the pilot reports or deciding for pilots how much equipment can fail or be unreliable before they ‘agree’ or decide to stop GPS jamming.” https://asrs.arc.nasa.gov/docs/rpsts/ctlr.pdf
jjwiseman
我运营着GPSJAM.org,过去五年一直在研究追踪GPS干扰对航空的影响,也是这篇报道的消息来源之一。我不是飞行员,也不为任何航空机构工作。
我非常期待看到美国国家运输安全委员会的最终报告,但初步报告和其他分析人士的评论似乎表明,这架机组的成员做出了错误的选择,并因此丧生。同时似乎也显示GPS干扰是导致事故的因素之一,如果美军没有干扰GPS信号,他们很可能还活着。
我认为飞行员需要做好在没有全球导航卫星系统的情况下飞行的准备,而军方定期干扰搭载着成千上万乘客的民用航班的GPS信号或许也是个糟糕的主意。
缺乏GPS并非关键安全问题,但确实是个安全问题。GPS干扰会减少应对选项,损害态势感知能力,导致其他安全设备(比如专门用来警告飞行员即将撞山的“地形感知与告警系统”)失效或性能下降,还会增加机组和空管人员的工作负荷与分心程度。
说“这种情况正是飞行员训练的内容,他们本应使用VOR和ILS导航”的人,暴露出一种过于大男子主义、对航空安全理解浅薄的态度。我们之所以增加GPS、ADS-B、TCAS、GPWS、TAWS等系统,是因为它们让飞行更安全。航空业和各国政府之所以对GPS干扰感到担忧,正是因为航空安全问题——如果你不担忧,说明你对问题的理解非常狭隘(“我的GPS如果失效了,我也绝不会遇到麻烦”)。
事实上他们认为这是个紧迫的问题。
2024年:“航空业寻求紧急解决方案应对GPS干扰” https://www.reuters.com/business/aerospace-defense/aviation-sector-seeks-urgent-solutions-gps-interference-2024-01-24/
来自2023年欧洲航行安全组织在联合国全球导航卫星系统/干扰检测与缓解国际委员会上的演示报告 https://rntfnd.org/wp-content/uploads/Aviation-GNSS-interference-UN-ICG-WGS-IDM-ECTL-GNSS-RFI-SEP23.pdf :
• GPS问题报告在所有类型的安全报告中占主导地位 • 冗余系统是航空业尽管面临全球导航卫星系统射频干扰却能维持正常运行的唯一原因! • 全球导航卫星系统已集成到众多系统中;精确的射频干扰影响难以预测,制造商不得不发布针对特定机型的指导文件 → 复杂性和工作量增加
航空安全影响 • 航空安全建立在两大原则上: • 信任你的仪表 • 遵守标准操作程序 • 全球导航卫星系统射频干扰导致飞行员不得不质疑这两条原则! • 一家大型航空公司的首席运营官: 导航不是我的问题。我的问题是“违规常态化”! • 已经发生过一些事件,仅仅是因为飞行员要处理太多系统警报而分心导致的。
美军干扰民用航班GPS信号的情况——在这次事故之前已经发生过几次险情,专家们多年来一直警告,如果军方继续这样做,即使发布了航行通告,也终将有人受伤,而这一预测不幸成为了现实。
来自2021年:“FAA文件揭示航空安全的一个惊人威胁:美军GPS测试” https://spectrum.ieee.org/faa-files-reveal-a-surprising-threat-to-airline-safety-the-us-militarys-gps-tests :
去年五月的一天凌晨,一架商业客机正在接近西得克萨斯州的埃尔帕索国际机场,驾驶舱内突然弹出警告:“GPS位置丢失”。飞行员联系了公司运行中心,得知新墨西哥州中南部的美国陆军白沙导弹靶场正在干扰GPS信号。“我们当时就知道,这不是飞机GPS的故障,”飞行员后来写道。
由于强风,该飞行员在一条跑道上进近失败,然后绕回来再次尝试。“我们被迫在黎明前使用04号跑道降落,且无法使用带垂直引导的仪表进近,”飞行员写道。“04号跑道……由于当地地形升高,存在很高的可控飞行撞地风险。”
此外
这远非涉及GPS干扰的最令人担忧的航空安全报告系统报告。2018年8月,爱达荷州一架在烟雾条件下飞行的客机据报告因军事测试遭受GPS干扰,若非空中交通管制员在最后时刻介入,飞机险些撞山。“生命损失可能发生,因为空管和机组相信他们的设备正常工作,但实际上却被引导着撞向山体,”该管制员写道。“如果(我们)没有注意到,那架机组和乘客就已经死了。我毫不怀疑。”
以下是一位拥有18年经验的空管人员在2024年关于白沙导弹靶场GPS干扰所写的内容:“如果无视飞行员报告,或者替飞行员决定在他们‘同意’或决定停止GPS干扰之前,有多少设备可以失效或不可靠,就会有人受伤。”https://asrs.arc.nasa.gov/docs/rpsts/ctlr.pdf
https://news.ycombinator.com/item?id=49187788
Wow. Jeff and Sanjay both departing. Truly end of a golden era.
There is an entire cohort of work-optional very senior engineers for whom one of the last reasons for hanging on was “at least Jeff and Sanjay are around”.
gandalfgeek
哇。杰夫和桑杰都离开了。一个黄金时代真正结束了。
有一整批可选择性工作的资深工程师,他们坚持留下的最后理由之一就是“至少还有杰夫和桑杰在”。
2026-08-05 08:29:31
- 大语言模型使通才任务普及,但领域专业知识仍是高效利用模型的关键瓶颈。
- 博客中AI生成图片会让人怀疑内容真实性,作者呼吁个人博客避免使用。
- Xbox宕机导致光盘游戏也无法游玩,凸显数字时代“拥有”物理媒体的虚幻。
- 《柔雨将至》讲述智能房屋在核战后徒劳运行,反思科技与人类命运的脆弱。
- 项目通过PCA定义肤色色彩空间,帮助开发者生成多样化肤色,但存在局限性。
- FFmpeg 9.0发布,新增动画WebP解码、GPU执行等更新并修复多项问题。
- 项目实现在单AMD MI300X上运行DeepSeek V4 Flash,性能达830 tok/s。
- 《柔雨将至》中自动化房屋在核废墟中徒劳运转,最终被大火摧毁。
- 美国在伊朗战争中几乎耗尽远程精确导弹库存,影响对其他国家的威慑能力。
- 苹果升级诉讼,指控更多前员工向OpenAI窃取机密数据,OpenAI否认指控。
https://www.seangoedecke.com/llms-reward-expertise/
在 2010 年代,如果你有技术短板(比如不会写 CSS),只能依赖有经验的同事或祈祷网上有现成答案。如今,通过将任务交给大语言模型,每个人都能写出还算过得去的 CSS,LLM 让所有人都成了通才。
因此很多人认为使用 LLM 不需要技巧——只要直接提问就能获得博士级数学、尚可的代码或 LinkedIn 风格的文字。但这是错误的:提示词最重要的技巧恰恰是你所提问领域的专业知识。
以数学家陶哲轩与 ChatGPT 关于雅可比猜想反例的对话为例。他的提示词简短精准,不逐条回应模型,只抓要点;模型输出也更简洁,因为陶哲轩通过展示专业度让模型进入"与数学家对话"模式而非"向业余者解释"模式;当模型回答有误时,他不直接反驳,而是说"这比我预想的复杂";他几乎从不采纳模型关于下一步的建议,而是自己提出跳跃性思路。
关键在于他真正理解数学——能从模型的多段回复中提取相关想法,提出替代方案,识别"看起来不对劲"的地方。这种"领域知识让你更善用 LLM"的理念同样适用于编程:如果你对代码库有良好理解,就能比不熟悉时更有效地驱动 LLM。
领域知识的价值表明,即使模型越来越强大,人类专业知识仍将持续有用。对许多任务而言,瓶颈是人而非模型,因为困难在于向模型精确传达人类想要的解决方案类型——信息"已在模型中",但需要非常聪明的人才能将其提取出来。
https://news.ycombinator.com/item?id=49161518
https://nelson.cloud/ai-generated-images-discourage-me-from-reading-your-blog/
作者在个人博客中表达了对 AI 生成图片的日益厌恶,尤其是在独立博客里看到这类图片时,会不禁怀疑文章内容是否也由 AI 生成。他认为个人博客出现 AI 图片令人失望,而企业博客出现尚可理解。作者表示,宁愿看到粗糙的微软画图作品,也不愿看到 AI 图片。他坦言自己的博客或许有很多可吐槽之处,但至少能确定是真实人类的思考,而非大语言模型的输出。最后,他呼吁个人博客作者避免使用 AI 生成图片。该文还附有 Hacker News 上的讨论链接。
https://news.ycombinator.com/item?id=49167113
https://birchtree.me/blog/xbox-goes-down-you-cant-play-games-you-own-on-disc/
Xbox 大规模宕机,导致玩家无法游玩自己拥有的光盘版游戏。文章引用 Jay Peters 的报道,指出此次宕机不仅影响数字游戏,连光盘游戏也被阻止运行。作者借此反思物理媒体的现状:如今的游戏光盘并非真正的“拥有”,游戏仍需安装并依赖网络验证,与过去 Game Boy 卡带即插即玩的体验截然不同。相比之下,PC 平台虽然也是数字版,但玩家有更多方式保持对游戏的访问,这也是作者转向 PC 的原因。
https://news.ycombinator.com/item?id=49167448
https://short-stories.co/@raybradbury/there-will-come-soft-rains-6k8vr4xxlnmj
在未来的 2026 年 8 月 4 日,加州艾伦代尔市的一座智能房屋仍在自动运行:它按时准备早餐、打扫卫生、播报天气,但主人早已在核战争中消失。房屋是废墟城市中唯一幸存的建筑,西墙上印着主人一家被瞬间烧灼的剪影。一只垂死的狗闯入后死去,被房屋自动处理。下午,房屋自动播放音乐、准备茶点,但无人享用。夜晚,房屋的语音系统为已故的女主人朗诵萨拉·蒂斯代尔的诗歌《柔雨将至》,诗中描绘自然在人类灭绝后依然生机勃勃的景象。午夜,一棵断树砸碎厨房窗户,引发火灾。房屋启动所有灭火系统,但最终因燃料耗尽而失控。在混乱中,房屋的各个功能模块仍在疯狂运转:烤面包机不断制作早餐、语音系统重复报时和朗诵诗歌。最后,房屋倒塌,只剩一面墙上的语音反复播报:“今天是 2026 年 8 月 5 日……”故事通过自动化房屋的徒劳挣扎,反思科技与人类命运的脆弱,以及自然对文明消亡的漠然。
https://news.ycombinator.com/item?id=49166491
https://toneyalexander.github.io/inclusive-color-space/
该项目旨在定义一个“足够好”的皮肤色调颜色空间,帮助开发者构建更具包容性的色彩工具(如角色创建器、数字艺术)。作者通过手动标记大量 RGB 颜色,利用主成分分析(PCA)将数据变换到更易处理的形状,并用数学方程将一个球体映射到该空间中,从而生成一组连续的、可覆盖广泛人类肤色的颜色范围。
页面提供了一个基于该颜色空间的交互式取色器,以及 Python 和 JavaScript 示例代码,方便直接使用这些数学公式。作者强调这些结果是“够用”的起点,而非权威标准,并坦诚讨论了局限性:肤色受血流、黑色素、散射等生物因素影响,且健康条件可能导致非典型肤色;个人主观偏差、屏幕和光线差异也会影响效果。最后,作者提醒技术处于社会背景中,浅肤色常被优先,而深肤色被边缘化,呼吁关注种族与肤色歧视问题。
https://news.ycombinator.com/item?id=49170165
https://github.com/FFmpeg/FFmpeg/blob/n9.0/RELEASE_NOTES
FFmpeg 项目发布 9.0 版本,代号“Lei”。该版本距离上一版 8.1 约四个月。完整更新日志位于项目根目录,Git 历史可访问 git.ffmpeg.org。如有问题,可通过 #ffmpeg IRC 频道(irc.libera.chat)或邮件列表联系。
https://news.ycombinator.com/item?id=49166202
https://github.com/ryanzhou/deepseek-v4-flash-mi300x
DeepSeek V4 Flash 在 AMD MI300X 上的生产级部署
项目概述 这是一个开源仓库,提供在单个 AMD MI300X GPU 上生产运行 deepseek-ai/DeepSeek-V4-Flash-0731 模型的完整配置和补丁方案,包含 Docker Compose 栈、SHA-256 校验文件覆盖、上游差异补丁及调优表格。
性能数据
为什么选择 MI300X MI300X 拥有 192 GB HBM3 和 5.3 TB/s 内存带宽,容量是 H100 SXM5 的 2.4 倍,价格约为一半。整个 304B 参数模型可完全放入 HBM,无需 PCIe 权重流式传输或分层卸载,单卡即可处理 2–8 个典型并发流及最多 64 流突发。
技术挑战与解决方案 MI300X(CDNA3)实现的是 AMD/Graphcore 的 fnuz 变体 E4M3 FP8 格式,而 MI325X 及更新产品使用 OCP 标准 FP8,内核若不区分会导致两倍的缩放域误差。官方 vLLM 配方主要针对 NVIDIA 和更新 AMD 硬件,本仓库补充了 FP8 格式修正、高并发 MoE 路由修复、因果投机验证、CPU-KV 同步修复及未调优核形状等方案。
仓库内容
https://news.ycombinator.com/item?id=49166386
第一篇《细雨将至》讲述了一座全自动化房屋在 2026 年 8 月 4 日这一天按部就班地运转:语音时钟报时、厨房自动烹饪早餐、机器人老鼠打扫房间、花园喷头洒水、晚间朗诵诗歌。然而这座位于加州艾伦代尔的房子是核战废墟中唯一幸存的建筑,屋主一家早已在原子弹爆炸中化为墙上的五道焦黑剪影。一只濒死的野狗误入屋内,最终死在客厅并被清理。当晚厨房突发火灾,房屋拼命自救却无力回天,最终在大火中坍塌。次日黎明,废墟中只剩一面断墙,墙上的电子语音仍在机械地重复着:今天是 2026 年 8 月 5 日。
第二篇《行人》设定在 2053 年的一个冬夜,作家伦纳德·米德每晚习惯在空无一人的街头散步,而全城三百万居民都闭门不出,沉迷于电视。一辆由电脑控制的警车拦住了他,用冰冷的金属嗓音反复盘问他的职业、住址和外出目的。当得知米德没有电视、没有妻子、以走路为乐时,系统判定这种行为属于"退行性倾向",将他押往精神病研究中心。警车驶过米德灯火通明的家,最终载着这位最后的行人消失在寂静的夜色中。
https://news.ycombinator.com/item?id=49162653
美国在伊朗战争中几乎耗尽了其远程精确导弹库存,包括陆军战术导弹系统(ATACMS)和精确打击导弹(PrSM)。这些武器对乌克兰战争也至关重要。分析人士担忧,库存下降可能限制美国威慑俄罗斯和中国等对手的能力。尽管特朗普声称美国拥有远超所需的弹药,且国防企业正以创纪录速度生产,但内部消息人士警告,持续冲突可能导致库存降至危险水平。此外,防御性武器如爱国者拦截弹和战斧巡航导弹的库存也大幅减少。
https://news.ycombinator.com/item?id=49166860
苹果公司在与 OpenAI 的商业机密诉讼中升级了法律行动,寻求法院发布初步禁令,阻止 OpenAI 利用苹果技术开发 AI 设备或产品。苹果在最新文件中指出,调查发现除已起诉的两名前员工(资深系统工程师刘畅和首席硬件官谭永毅)外,还有 11 名前苹果员工可能涉及窃密,包括有人曾在面试 OpenAI 前与被告会面讨论苹果未发布产品信息,还有人截取机密文档截图。苹果要求法院允许加速证据开示。OpenAI 公开回应称苹果的禁令请求“基于虚假信息且完全没必要”,强调自身没有也不想获取苹果商业机密,并指出苹果此前曾因混淆相似姓氏而发错邮件、未如实说明与法律顾问的沟通等失误。
https://news.ycombinator.com/item?id=49170479
https://news.ycombinator.com/item?id=49163331
I did a test a few months ago. A friend of mine wanted to develop what i understood to be a simple single page web app. But since she didn’t have any software engineering experience she asked me to help. Around that time everyone was talking about how literally anyone can develop software with LLMs i asked her if she could give it a try first, and if I could watch the attempt.
I was fully expecting that writing the code will pose no problem for the AI. But i was curious if the AI will realise that my friend is a novice and needs extra help with things like: copy pasting the code into a text file and saving it with an html extension, helping her host the file online so she can share it with others, buying a domain for it, etc. I assumed they will get there eventually, but i also assumed that it will take a lot of stumbling around and misunderstandings.
But i was completely wrong. They didn’t even get to that point. Because my friend didn’t have the vocabulary to ask the AI to write code. They were just going around in circles where the AI was brainstorming with her about possible features and getting thints more and more complicated. We terminated the experiment after one and a half hours and many many messages exchanged between her and the LLM.
Whereas to me who knows the terminology would have probably just taken a single exchange of messages to get the result she described to me. I would have prompted with something like “Please write to me an html page which does X, Y, Z.” But since she didn’t know the right terminology she got into a vortex of feature discusion, and she didn’t find a way to tip the AI into “just do it, write it now” mode. In other words in that case the LLM would have rewarded even just a little bit of expertise, but without it there was a confusion about goals between the human and the machine.
krisoft
几个月前我做了一个测试。我有个朋友想开发一个在我看来很简单的单页网页应用,但她没有任何软件工程经验,就找我帮忙。当时大家都在说,任何一个人都能用LLM开发软件,于是我问她能不能先自己试试,让我在旁边看着。
我原本以为,写代码对AI来说完全不是问题。但我好奇的是,AI会不会意识到我朋友是个新手,需要额外帮助,比如:把代码复制粘贴到文本文件里,保存成html扩展名;帮她托管文件以便分享给他人;以及购买域名等等。我猜他们最终能搞定,但过程中肯定会经历很多磕磕绊绊和误解。
但我完全猜错了。他们甚至都没走到那一步。因为我朋友不知道该怎么跟AI说要写代码所需的关键词。他们一直在兜圈子,AI跟她头脑风暴各种功能,把需求越搞越复杂。我旁观了一个半小时,她跟LLM来回发了好多条消息,最后我们终止了实验。
而对我来说,只要懂术语,可能只需要一条消息就能得到她描述的那个结果。我会直接提示:“请写一个实现X、Y、Z功能的HTML页面。”但她不知道正确术语,就陷入了功能讨论的漩涡,找不到办法让AI切换到“别废话,直接写”的模式。换句话说,在这个案例中,LLM哪怕奖励一点点专业知识就能高效工作,但缺乏专业知识时,人和机器之间就陷入了目标混乱。
https://news.ycombinator.com/item?id=49169001
I wanted to show my gf the halo 1 campaign, so I downloaded the 30GB Masterchief Collection via steam. When I launched it, I was greeted with a black box. After 1 Minute it was a microsoft login screen that looked eerily like a browser inside the game. Since I couldn’t do anything else, the game was stuck at a sub 720p resolution. I said fine Ill just create an account. Yea so mail, firstname, lastname, birthday, country, email and email verification code later I had to do captcha which is clicking and holding down a button. “Help us against the bots” it said. I unironically spent 5 minutes clicking that button, then being stuck in a loop with “please try again” while my gf just sat there watching me.
Well I turned off MCC and gave up.
I don’t want to know what life is like having xbox hardware
mawadev
我想给女朋友展示一下《光环1》的战役,于是通过Steam下载了30GB的《士官长合集》。启动游戏后,屏幕先是一片黑,过了一分钟才出现一个微软登录界面,看起来诡异得像游戏里嵌了个浏览器。我什么也做不了,游戏还被锁定在不到720p的分辨率。算了,我心想,那就注册个账号吧。于是填了邮箱、名、姓、生日、国家、邮箱验证码,最后还得做一个按住按钮的验证码。上面写着“帮助我们对抗机器人”。我老老实实花了一分钟按那个按钮,结果陷入“请重试”的死循环,女朋友就坐在旁边看着我。
好吧,我关掉了《士官长合集》,放弃了。
我都不敢想用Xbox硬件会是怎样的体验。
https://news.ycombinator.com/item?id=49164936
Is it normal for a company this big to make a public blog post like this with almost no introduction to what they’re on about?
I would expect them to start out with something like Apple is suing us for yada yada yada, but instead it reads like the diary of a hurt teenager.
efnx
像这样一家大公司,发表一篇几乎没介绍自己到底在说什么的公开博客文章,这正常吗?
我本以为他们会先写点类似“苹果正在起诉我们,因为blah blah blah”的话,但这读起来却像一个受伤青少年的日记。
https://news.ycombinator.com/item?id=49164268
I know everyone wants to crap all over these setups that are impractical, but this is how progress happens.
People will keep plugging away at this and figure out how to avoid wearing the hard drive, how to make it run faster, custom hardware buses etc.
Keep going! I personally can’t wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.
dghlsakjg
我知道大家都会吐槽这些不实用的配置,但进步就是这样发生的。
人们会不断钻研,想办法解决如何避免佩戴硬盘、如何让它运行更快、定制硬件总线等问题。
继续加油!我个人已经迫不及待想看到那一天:一个1万亿参数的模型能在200美元的SSD上运行,而不是靠5万美元的Nvidia芯片机架。
https://news.ycombinator.com/item?id=49169079
This is my experience with all Microsoft software. The worst thought out, buggy, untested pile of crap.
There’s something seriously wrong in Redmond.
At least with linux, you can search on the web and find a forum post telling you the magic incantation. With Microsoft, you will be forced to consume some second-rate engineer’s unmanaged, barely workable, edge case until it might or might not work.
Office on the web is a horror show this way.
epistasis
这就是我对所有微软软件的体验:构思最差、漏洞百出、未经测试的一堆垃圾。
雷德蒙德绝对出了大问题。
至少用Linux,你可以在网上搜索,找到论坛帖子告诉你神奇的咒语。而用微软,你只能被迫忍受某个二流工程师编写的、无人维护、几乎无法正常运行的边缘案例,直到它可能生效,也可能不生效。
Office网页版在这方面简直就是一场噩梦。
https://news.ycombinator.com/item?id=49152482
A very nice read and pretty much my experience. I have been here for a quarter of my life now, I do not think I can live anywhere else now. People may not like me being here because of my skin color or whatever but I do not mind. I keep to myself and follow all the “rules” I can. In fact, I like most of the rules, rules are good because they ensure order, fairness and equality.
Will a little bit of empathy be nicer, sure, but I cannot complain. Speaking of which, a lot of German natives complain and some are even leaving the country (around 100k left last year), but honestly I have visited / lived in nine countries, and people have no idea how good they have it here. I came here with a bag, twenty euros in my pocket and a job offer, a dozen years later I have pretty much all basic comforts and stability one can dream of, my kids have a bright future waiting for them, all this is not possible in many countries.
rockyj
非常棒的阅读体验,几乎就是我的亲身经历。我生命中的四分之一时光都在这里度过,现在我觉得自己已经无法在其他地方生活了。有些人可能因为我的肤色或其他原因而不喜欢我在这里,但我不介意。我独来独往,遵守所有我能做到的“规则”。事实上,我喜欢大部分规则,规则是好的,因为它们确保了秩序、公平和平等。
多一点同理心当然更好,但我没什么可抱怨的。说到这个,很多德国本地人却在抱怨,甚至有些人正在离开这个国家(去年大约有10万人离开),但说实话,我去过/住过九个不同的国家,人们根本不知道自己在这里过得有多好。我当初只带着一个包、兜里揣着20欧元和一份工作offer来到这里,十几年后,我几乎拥有了一个人能梦想的所有基本舒适与稳定,我的孩子们未来光明,而这些在世界上许多国家都是不可能实现的。
https://news.ycombinator.com/item?id=49173038
This is a very different problem than actual efficiency when mowing a lawn or vacuuming a rub. Turning takes more work and time, and you miss a part of the ‘square’ when you are turning (the arc of the curve is not a straight line). In addition, vacuums especially have an area at the edge of the machine that is not cleaned as well, which is why you want a bit of overlap in your lines.
Plus, when mowing a lawn, you often want to leave a nice pattern, so you cant just go for pure efficiency.
cortesoft
这与割草或吸尘时的实际效率截然不同。转弯需要更多精力和时间,而且在转弯时你会错过“正方形”的一部分(弧线并非直线)。此外,尤其是吸尘器,机器边缘区域清洁效果较差,因此你需要让清扫路线稍有重叠。
另外,割草时你通常想留下漂亮的图案,所以不能只追求纯粹效率。
https://news.ycombinator.com/item?id=49172861
I believe “money good” is YC’s only guiding principle. HN is a nice resource, but let’s not have any illusions about the morals of the owners.
chuckadams
我认为"金钱至上"是YC唯一的指导原则。HN是个不错的资源,但我们不要对拥有者的道德抱有幻想。
https://news.ycombinator.com/item?id=49171927
Apple is cutthroat in business too.
Tony Fadell, the inventor of the iPod and co-inventor of the iPhone, and later Nest founder, commented this in Stratechery about the lawsuit when filed:
“This is Apple’s typical tactic to scare Apple employees — either former or current. I heard this lawsuit was driven by the Apple board.
Steve threatened to file a lawsuit against Nest for poaching 80-100 Apple employees. He called me, screamed for a while with lots of accusations. Then I said, “Steve, it’s Apple’s job to retain its talent, not mine.” He stopped his rant and then we went on to talk about our families and vacation plans. We kept hiring…”
jdross
苹果在商业上同样冷酷无情。
iPod的发明者、iPhone的共同发明者、后来Nest的创始人托尼·法德尔在诉讼提起时于Stratechery上评论道:
“这是苹果恐吓苹果员工——无论是前员工还是现员工——的典型手段。我听说这起诉讼是由苹果董事会推动的。
史蒂夫曾威胁要起诉Nest,因为我们挖走了80到100名苹果员工。他给我打电话,大吼大叫了一阵,满是指责。然后我说:‘史蒂夫,留住人才是苹果的事,不是我的事。’他停止了咆哮,然后我们转而聊起了彼此的家庭和度假计划。我们继续招人……”
https://news.ycombinator.com/item?id=49167629
Gaming is heading in the same direction that TV, movies and music are. It’s a shame. I can invite my friends over, boot up my GameCube to play some Mario Kart Double Dash. As long as the console and disc are okay, I can (hopefully) do that until I pass away.
These days? Not so much. I’ll never be able to be sure I can play GTA VI in twenty years. It just sucks not owning anything anymore. I know why they do it, it’s a massive profit difference.[1]
On the PC end of things, it’s not as bleak. Services like GoG let you download the standalone installer, DRM free. I try to get any single player games I can from GoG, you should too.
Support the corps that give you ownership. Buy those Blu-rays of your favorite films. Either to support physical media or to ensure you will have it until data degrades.
1: https://youtube.com/watch?v=sESfDf0My84
cautiouscat
游戏正朝着电视、电影和音乐相同的方向发展。真是可惜。我可以邀请朋友来家里,打开我的GameCube玩《马里奥赛车 双重冲击》。只要主机和光盘没问题,我(希望)能一直玩到老死。
现在呢?没那么简单了。我永远无法保证二十年后还能玩到《GTA VI》。什么都再也无法真正拥有,这感觉糟透了。我知道他们为什么这么做,利润差距实在太大了。[1]
在PC端,情况还没那么糟。像GoG这样的服务允许你下载独立安装程序,没有DRM限制。我尽量从GoG购买所有能买到的单人游戏,你也应该试试。
支持那些让你真正拥有产品的公司。去买你最爱电影的正版蓝光碟吧——要么是为了支持实体介质,要么是为了确保你能一直拥有它,直到数据自然降解的那一天。
1: https://youtube.com/watch?v=sESfDf0My84
https://news.ycombinator.com/item?id=49162323
The amplifying mirror analogy works best here. LLMs are ultimately a reflection of your own interactions with its weights, the tone you use, the structure with which you construct your prompt, aspects of an issue you tend to focus on, your breadth of vocabulary and world knowledge and whatnot.
People who (carefully) use it as an extension of their own mind and senses will very likely thrive, and those who use it as a replacement for their minds and their senses will struggle.
One of the Claude skills I made Claude itself generate was the ’learning a concept across tiers’ skill – from ELI5 level to a PhD level, and it triggers whenever I ask it a very general question on a complex topic that isn’t my bread-and-butter. The fact that I’m able to choose explanation level from a super smart LLM (that’s available 24x7) that can explain any topic under the sun would’ve been mind-bogglingly sci-fi-ish just 4 years ago in 2022.
abixb
放大镜的比喻在这里最为贴切。LLM本质上是你与自身权重互动的映射——你使用的语气、构建提示词的结构、你倾向于关注问题的哪些方面、你的词汇量与知识广度等等。
那些(谨慎地)将其作为自身思维与感官延伸的人,很可能会蓬勃发展;而那些将其作为思维与感官替代品的人,则会举步维艰。
我让Claude自行生成的一项技能是“跨层级理解概念”——从五岁孩童能懂的水平到博士级别,每当我提出某个我不擅长的复杂领域中的宽泛问题时,它就会被触发。我能够从一个超级智能的LLM(且全天候可用)那里选择任意议题的解释层级——就在四年前的2022年,这还绝对属于令人难以置信的科幻情节。
https://news.ycombinator.com/item?id=49172620
Police use the technology to investigate crimes such as hit-and-runs and to locate missing persons. > No suspects have been identified.
Yeah, that tracks.
Camera data is owned by the police department and permanently deleted after 30 days.
I assume this means that the police department deletes one local copy?
jeremyberemy
警察利用该技术调查肇事逃逸等犯罪案件,并定位失踪人员。>尚未锁定任何嫌疑人。
确实,这样说得通。
摄像头数据归警察局所有,30天后永久删除。
我猜测这意味着警察局删除了一个本地副本?
https://news.ycombinator.com/item?id=49164840
Setting aside the presented evidence, it feels weird for somewhat emotionally charged complaining like this to go on an official company blog. Companies are generally pretty tight-lipped about active litigation outside court documents, right? This just seems a bit amateurish, down to the title. Is public opinion about this case so important to them?
gr_norm
撇开已呈现的证据不谈,这种情绪化的抱怨出现在公司官方博客上感觉很奇怪。公司通常对除法庭文件外的未决诉讼都闭口不谈,对吧?这种做法从标题到内容都显得有些业余。这个案件的舆论导向对他们来说就这么重要吗?
https://news.ycombinator.com/item?id=49174272
I live very close to LAX, and we have hundreds of Waymos in our area. At first i was a little weary, and kinda weirded out by all of them.
They have become completely normal, and cause WAY fewer traffic incidents than human drivers. They are very predictable, and will always let you pull in front of them when changing lanes (unlike a lot of other LA drivers).
I have seen a few incidents where they seem ‘stuck’, and it was a bit annoying and kind of funny, but again, it does not happen nearly as often as for human drivers.
I also like taking them. It is nice to just get in and relax and not have to talk to anyone or worry about who is going to pick you up. You can control the temperature and music from your phone, you can drop the pin exactly where you want to be dropped off and picked up, and the ride is smooth and comfortable.
We can debate all the knock on effects of Waymos, but the user experience is pretty great.
cortesoft
我住在离洛杉矶国际机场很近的地方,我们这边有好几百辆Waymo无人车。一开始我还有点心里发毛,看到它们总觉得怪怪的。
现在它们已经完全成了常态,而且引发交通事故的概率比人类司机低得多。它们非常守规矩,你变道的时候它们一定会让你插进来(不像很多其他洛杉矶司机)。
我也见过几次它们好像"卡住"了的情况,虽然有点烦人但也挺好笑,不过话说回来,这种状况发生的频率比人类司机可低多了。
我也挺喜欢坐它们的。上车就能放松,不用跟任何人说话,也不用担心谁来接你。温度、音乐都能在手机上调节,上下车的地点也能精确设定,而且行驶过程又平稳又舒适。
我们可以讨论Waymo的各种连锁影响,但就用户体验而言,确实相当出色。
https://news.ycombinator.com/item?id=49173886
For the most-dominant military power in the world to “whoopsy” run out of weapons because of a low-stakes engagement isn’t readily believable.
Wow, yeah, you’d need to have morons at the top of the chain of command and subordinates afraid to speak out in order for something like that to happen. Nothing like the US government today at all.
justin66
世界上最强大的军事力量,竟然因为一场低风险冲突就“哎呀”一声把武器打光了,这实在让人难以相信。
哇,是啊,那得是指挥链顶端全是蠢货、下属又不敢发声,才会发生这种事吧。跟今天的美国政府可一点都不像呢。
https://news.ycombinator.com/item?id=49169265
It is strange to me that people are re-discovering this today, in 2026. Microsoft has always been half-baked, even in the 80s and 90s. Heck, my parents basically raised me (20+ years ago) by running a business fixing problems with Microsoft products for their clients.
keiferski
真奇怪,到了2026年人们才重新发现这一点。微软从来都是半吊子,哪怕在80年代和90年代也是如此。甚至,我父母20多年前基本就是靠为客户解决微软产品的问题来维持生计的。
https://news.ycombinator.com/item?id=49167897
AI-generated images discourage me from interacting with any entity AT ALL. Bar / restaurant posters, events, shops, anything. Because it makes the brand appear lazy. And in a weird way “anti-human”. I don’t have a better word than that.
wateralien
AI生成的图像让我完全不想与任何实体互动。酒吧、餐厅的海报、活动、商店,任何东西都如此。因为这会让品牌显得很懒惰,而且以一种奇怪的方式显得“反人类”。我找不到比这更合适的词了。
2026-08-04 07:04:30
- 协作中直接转发AI长篇内容而不消化是“肉代理”行为,主张必须自己验证并以自身语言表达才算贡献。
- Qwen3.8-Max发布,2.4万亿参数,在编码、研究复现等方面全面升级,能独立完成多日项目并开源权重。
- JFrog发现一批针对SQLite的“严重”CVE漏洞公告是AI生成的虚假内容,提醒不能仅凭描述采信漏洞情报。
- 文章认为AI降低代码修改门槛,用户可通过指令fork修改开源项目,但需重构产品以适应这种流动式生态。
- 作者作为土耳其移民在德国感受到友善平等的文化,自认比许多德国人更德国。
- OpenAI用Astra模型在数学与理论计算机科学领域取得十项新进展,并用Lean完成形式化验证。
- 作者坚持手动重敲AI生成代码以维持对代码库的完整理解,避免团队积累认知债务。
- Isopolis是用等距像素风格绘制的旧金山互动地图,展示科技地标并允许用户提交地点。
- 2025年德国风能与太阳能发电首次超越化石燃料,标志能源转型里程碑,但退煤仍面临挑战。
- Bonsai是Jane Street用OCaml构建的响应式Web UI库,内部广泛使用,实现前后端同类型系统。
https://gruhn.me/blog/2026-08-03/
作者批评了一种常见行为:在 Slack、代码评审或聊天中,直接转发 AI(如 Claude)生成的长篇回答,自己却不加消化。这种行为没有增加价值,因为对方本可以直接与 AI 对话,反而需要额外阅读冗长且可能包含“貌似合理但错误”的内容,甚至充满难懂术语。
作者主张:可以借助 AI 辅助思考,但不应机械转述输出。正确做法是主动阅读、理解并验证 AI 的结果,然后用你自己的语言重新表达——这样才能体现你的真正贡献。
针对代码评审场景,作者指出:很多人几乎不读代码,只是把任务描述和评审反馈原样粘贴给自动编程工具,再转发结果。这种情况下,真正的实施者其实是评审者(借助 AI),而操作者只是“肉代理”(meat proxy),毫无价值。
https://news.ycombinator.com/item?id=49151933
https://qwen.ai/blog?id=qwen3.8
今天,我们正式发布了 Qwen 3.8-Max,这是迄今为止 Qwen 系列中最强大的模型。这也是我们首次开源 Qwen-Max 类模型的权重,相关权重将在下周发布。Qwen 3.8-Max 基于 Qwen 3.5 的架构,参数扩展至 2.4 万亿,在编码、工作、研究和长时间任务方面提供了全面的改进。它不仅能够回答更具挑战性的问题,还能更可靠地完成复杂的任务,生成可靠的成果。
在编码方面,Qwen 3.8-Max 的能力远超简单的函数请求,它能够从空文件夹开始,独立完成真实的多日项目。我们对其进行了三项挑战测试,Qwen 3.8-Max 在没有人类干预的情况下,通过反馈循环自我演变,成功完成了所有任务。其中一个案例是构建了一个自我进化的工具(oh-my-cli 项目),该工具在超过 10 天的自主编码过程中,通过用户反馈和社区实践,不断迭代完善。Qwen 3.8-Max 在 GitHub 上公开了整个项目的跟踪记录。
在研究复现方面,我们让 Qwen 3.8-Max 复现了一篇名为 “统一数据选择以进行 LLM 推理” 的研究论文。它从头开始设计和编写数据处理脚本、训练代码和评估设置,完成了约 7600 行代码,并进行了 33 轮 GPU 训练。最终,Qwen 38-Max 不仅复现了论文的实验结果,还在此基础上进行了改进,提出了 18 个改进思路。
我们还将 Qwen 3.8-Max 投入了真实的在线竞赛中。在 WWW2025 多模态对话意图识别挑战中,它在 24 小时内独立完成了全套解决方案,使用了多种语言模型和视觉语言模型,最终击败了 87% 的参赛人类团队。
在工作方面,Qwen 3.8-Max 的能力表现也十分出色。它在多个高经济价值的职业中表现出色,例如在公司合规审查中,能在一个小时内完成对数百份文件的审查,通常需要一周的时间。而在 UI/UX 设计方面,它能一次性生成高保真交互原型,而不需要人类的修订。此外,Qwen 3.8-Max 还在餐饮业、结构工程、康复治疗等多个领域展现了其工作能力,显著提升了人类的工作效率。
Qwen 3.8-Max 的动态工作流功能使其能够编排大规模的子代理系统,从而实现复杂的任务规划和执行。它在量化研究方面的能力也引人注目,能够并行化处理广泛的研究方向,将传统的数周工作压缩至单次会话内完成。
综上所述,Qwen 3.8-Max 通过其强大的编码、研究、工作能力和动态工作流,展示了在长时间任务上的潜力,标志着人工智能在各个领域应用的进一步深化。
https://news.ycombinator.com/item?id=49150470
https://research.jfrog.com/post/sqlite-critical-cves-or-llm-slops/
JFrog 安全研究人员发现,一批针对 SQLite 的“严重”CVE 漏洞公告极有可能是由 AI 生成的“LLM slop”(低质虚假内容)。这些公告来自一个新创建的 GitHub 仓库,并被 NVD 标记为严重级别。
经核查,这些漏洞描述存在多处硬伤:引用的函数在目标版本中根本不存在,PoC 无法触发崩溃,修复补丁无处可寻,且均未出现在 SQLite 官方公告中。例如,CVE-2026-51302 提到的函数在 3.41 版本中尚未引入;CVE-2026-51303 声称的修复版本差异中并无相关改动;CVE-2026-51296 引用的行号甚至超出了源文件总长度。
研究人员在隔离环境中用 ASan 编译了官方 SQLite 并执行 PoC,均未复现任何内存错误。这些虚假 CVE 之所以能通过审核,是因为 MITRE 提交系统缺乏身份验证,而 NVD 自 2024 年 2 月起因报告激增暂停了深度人工分析,导致伪造公告钻了空子。此次事件再次敲响警钟:当前漏洞生态中,仅凭“看起来合理”的描述已不足以采信。
https://news.ycombinator.com/item?id=49154332
https://blog.exe.dev/devtools-must-be-open-source
这篇文章讨论了 AI 代理时代下,软件开发个性化方式的根本转变。作者认为,过去为个人定制软件成本高昂,如今借助 AI 代理,只需简单指令即可下载源码、本地构建并自动同步上游更新,使得持续维护个人修改变得容易。文章以作者的个人项目 meat.dev 为例,展示了如何通过一条提示词将其集成到 Shelley 代理中,实现自动预处理代码审查。核心观点是:AI 代理降低了定制软件的门槛,传统软件依赖配置文件和插件系统的模式将被“直接修改源码”取代,许多软件产品类别需要为此重新设计。
https://news.ycombinator.com/item?id=49156111
https://mertbulan.com/more-german-than-many-germans/
作者是来自土耳其的计算机专业学生,2017 年通过 Erasmus 奖学金到德国汉堡实习,原计划只是为毕业后求职积累经验,但这段经历彻底改变了他的人生。
初到德国时,他打破了此前对德国人的刻板印象。他的德国室友友善耐心,对他表现出极大的信任;工作团队友好热情,带他参与团队活动,给他超出普通实习生的责任。整个夏天成为他人生中最美好的时光,公司也直接给了他全职工作邀请。
2018 年毕业后,他正式搬到汉堡工作生活。从注册住址时工作人员的一句“欢迎回来”,到同事们帮他找房、解决电费问题、搬家,他感受到家一般的温暖。德国职场文化也让他印象深刻:管理层平易近人,同事关系平等,尊重每个人的饮食需求,强调休息和信任。
他注意到德国社会阶层差异小,建筑工人和白领在同等价位的餐厅用餐,富人区物价也与其他区域相差无几,这让他感受到社会民主的真实存在。
在工作中,他没有感受到移民歧视,反而当选为工会委员中得票最高的人。他也提到拥有土耳其血统的政界人物在德国身居高位,进一步印证了德国社会的包容与平等。
最终他得出结论:现在的他“比许多德国人还要德国”,这段经历让他选择留在这个国家,并真正认同了这里的社会价值。
https://news.ycombinator.com/item?id=49151734
https://openai.com/index/ten-advances-in-mathematics/
OpenAI 发布了一项重要成果:在数学与理论计算机科学领域取得了十项新进展,均由内部版 Astra 模型生成,并用 Lean 形式化验证,人类参与整理成论文。这些结果解决或大幅推进了多个长期未决的公开问题,涵盖高维几何、编码理论、电路复杂度、群论、算子代数、量子复杂度、格密码和极值组合等方向。
这十项成果包括:高维球堆积的新上界;二进制码和球面码最大尺寸的指数级改进;非 sofic 群存在性的构造证明;推翻了 Connes 刚性猜想;永久项计算的算术电路下界;量子二人博弈的指数并行重复定理;最近向量问题多项式因子不可近似;Ehrhart 体积猜想各维度下的完整解答;multicolor Ramsey 数的超指数下界;以及极值图论中紧致性与退化性猜想的进展。
OpenAI 同时宣布为十万名科学家和数学家提供免费 ChatGPT 使用权限,并强调会诚实标注 AI 在研究中的贡献,希望数学社区深入审视和推进这些结果。
https://news.ycombinator.com/item?id=49157930
https://ankursethi.com/blog/prevent-cognitive-debt-by-manually-retyping-llm-generated-code/
作者在自己的个人项目中仍使用 AI 编码助手,但不再允许它直接修改代码。他发现自己无法理解 AI 生成的代码,也不想审查那些冗长且质量可疑的 AI 输出,于是采取了一种“低效但清醒”的方法:让编码助手把建议的代码和命令显示在聊天窗口里,由他手动输入到项目文件中。
他在代理配置中加入了明确指令:不得创建、编辑或删除文件,也不得运行修改项目的命令,只能展示建议。这样他虽然只达到两倍开发速度,但每行代码都过了一遍大脑,能及时察觉幻觉或糟糕设计,并顺手重构、注释,形成对代码库的完整空间认知。
作者类比年轻时学习编程的“手动抄写”方式,认为这是在 AI 时代保持理解力的工作流。他担心整个软件行业正积累认知债务,而他至少要确保自己发布的软件是真正理解的。
https://news.ycombinator.com/item?id=49153374
Isopolis 是一个以等距风格绘制的旧金山互动地图网站,融合了硅谷文化与城市导览功能。页面支持浏览街区、地标、公司(如 Airbnb、Anthropic、OpenAI、Y Combinator 等),并提供了多条特色旅行路线,包括经典旧金山入门游、初创公司主题游、Twitter 现实游和徒步挑战游。用户可在地图上查看事件、标记地点,并通过社区贡献功能提交新地点(需填写名称、类型、坐标、说明等)。网站还播放类似“硅谷主题”的背景音乐,设计风格致敬 isometric.nyc,由 Nuwan Davek 创建,描述文本部分由 AI 生成。
https://news.ycombinator.com/item?id=49149966
2025 年,德国的风能和太阳能发电首次超过化石燃料,标志着一个重要的里程碑。根据 Carbon Brief 对《世界能源统计回顾》数据的分析,风能和太阳能合计发电量达到 225 太瓦时(TWh),占总发电量的 44%,而化石燃料的发电量为 217 TWh,占 43%。这一变化反映出德国在 “能源转型”(Energiewende)战略下,经过 20 年的快速增长,正逐步摆脱煤炭和核能。
德国的目标是到 2030 年安装 115 吉瓦(GW)的陆上风能,并在 2025 年批准了 20800 兆瓦的新装机容量。官方目标要求到 2045 年实现经济全面净零排放,到 2030 年电力消费中可再生能源占比达到 80%,并希望到 2035 年基本实现气候中性电力系统。
由于核能的逐步淘汰,德国在实现这些目标时比邻国如法国和英国更依赖可再生能源。核能的淘汰是 “能源转型” 的核心内容,管最近有政治上的反对声音,但这一政策依然得到广泛认可。今年早些时候,右翼中间派总理弗里德里希・梅尔茨称核能淘汰是一个 “战略错误”,但他的政府排除了恢复传统核能的可能性。
煤炭依然是德国面临的更大短期挑战。德国的煤炭使用量仍远高于大多数其他欧洲国家,官方的煤炭淘汰截止日期为 “最迟在” 2038 年。然而,专家认为,德国在淘汰煤炭方面的进展可能会在这一日期之前几年完成,尽管在近期能源危机期间出现了放慢转型的压力。
当前,可再生能源面临另一种反对声音,来自于右翼政党 “德国选择党”(AfD)。与此同时,现任联盟政府也在推进新燃气发电厂的建设,这些发电厂被立法视为过渡技术,计划到 2045 年转换为绿色氢气发电,以保持与气候中立目标的一致性。尽管几乎没有其他声音主张完全放弃煤炭淘汰,但政府预计将在 8 月发布对其时间表的审查,这将是考验柏林是否坚定支持转型政策的下一个重要时刻。
https://news.ycombinator.com/item?id=49155359
https://github.com/janestreet/bonsai
Bonsai 是一个用 OCaml 构建高性能响应式 Web 应用的 UI 库,部分灵感来自 Elm,被 Jane Street 内部几乎所有 Web 应用广泛使用。
组件采用纯函数式状态机实现,易于组合;框架内部的增量机制确保仅在相关状态变化时才重新计算,适用于所有值,而不只是视图。
Bonsai 将状态、增量性和渲染解耦,可按需组合;状态管理不绑定具体组件,支持复杂的生命周期和局部状态托管。同时,使用 OCaml 可实现前后端共享类型和业务逻辑,增强代码可维护性。
它还提供强大的模板语言、组件级样式表,以及自动化测试系统,可通过编程操作 UI 元素并观察 DOM 变化,编写表达力强的期望测试。
https://news.ycombinator.com/item?id=49152842
https://news.ycombinator.com/item?id=49152018
I deal with this all day long at work and it’s exhausting. People almost acting like no one has thought of it “I asked Claude what happened, and it spit out this 300 line response. Can you read it for me and see if it’s right?”
What kills me is you might expect this from a busy high level manager that doesn’t really understand the technical details and they just point the AI to an error they got. They don’t know how to interpret the response, so they ask someone who work on the thing. It’s still kinds annoying because you could just ask, but whatever. But to get these from junior and senior engineer for the areas they work in and expect someone else to read it for them? It’s crazy behavior. How can someone serious even think that’s ok.
eddythompson80
我整天在工作中处理这种事,简直累死了。人们表现得好像没人想到过似的——“我问了Claude发生了什么,它吐出了一段300行的回复。你能帮我看一下对不对吗?”
最让我崩溃的是,你也许以为这种事儿只会发生在那些不了解技术细节的忙碌高管身上,他们只是把遇到的错误丢给AI,不知道怎么解读回复,就去找懂行的人问。这虽然也挺烦人的,因为你本来可以直接问,但算了。可是,如果是初级甚至高级工程师在自己负责的领域里也这样,还指望别人替他们读回复?这简直疯了。一个认真做事的人怎么能觉得这样没问题?
https://news.ycombinator.com/item?id=49150809
They’ve also announced Qwen3.8-27B being released open-weight next week. Qwen3.6-27B is widely regarded as one of the best local models, especially since nothing else comes close to it, that isn’t benchmaxxed, without being significantly larger. If 3.8 truly improves upon it that would be awesome.
toshinoriyagi
他们还宣布下周将开源发布Qwen3.8-27B的权重。Qwen3.6-27B被广泛认为是最好的本地模型之一,尤其是因为其他接近它的模型,要么是跑分特化型,要么体积明显更大。如果3.8真的能在此基础上有所改进,那将非常棒。
https://news.ycombinator.com/item?id=49152248
At my dayjob there is a person spearheading ai across the enterprise.
They generated lots of documentation across the whole stack and now makes all PO/BAs read it if it’s correct. So not just 300 lines - he unironically generated thousands of lines of “documentation” and is now making hundreds of people review it for him
Complete brainrot
Au psychosis is getting seriously outrageous at this point
Thankfully I’m a dev and thus aren’t in the blast radius of that genius idea
ffsm8
在我的日常工作里,有个人正在全公司牵头推行AI。
他生成了涵盖整个技术栈的大量文档,现在让所有产品经理和业务分析师去读,检查是否准确。所以不只是三百行——他真就搞出了几千行的“文档”,然后让几百号人替他审阅。
完全是脑残行为。
到这一步,Au 精神病真是越来越离谱了。
还好我是开发人员,所以不在那个天才主意的影响范围内。
https://news.ycombinator.com/item?id=49150151
Oh, this again. I should put a website with this up…
I was the person who personally ran 10.6 security updates at Apple (10.6.1+), the “DRI”. My team in the Updates Program office and I reviewed every single bug to determine if it should go in a security and stability update or wait for the next major version. Seriously, every morning we group triaged all Mac OS X bugs, both incoming and those nominated internally for us to look at and determine if it should go in an update. I packaged and audited the builds and tuned the delta vs full updates. I built the system that largely automated diffing “trains” for software updates (automastering).
The new version of the OS was always being developed in a branch/train, and fixes were backported to the current version as they were found. They weren’t developed linearly / one after another. So, if you are comparing the most stable polished/fixed/stagnant last major version with the brand new 1.0 major version branch, the newer major is going to be buggier. That would be the case with every y.0 vs x.8. But if you are comparing major OS versions, Snow Leopard was different.
Snow Leopard’s stated goal internally was reducing bugs and increasing quality. That is a fact, not marketing. I am not sure why people on the internet don’t believe that, but I was there. If you wanted to ship a feature you had to get explicit approval from leadership and the bar was high. In normal feature releases it operated bottom up “here is what we are planning to ship” and in Snow Leopard it was top down “can we ship this?”.
AFAIK Snow Leopard was the first release of this kind (the first release I worked on was Jaguar or Puma), and was a direct response to taking 8 software updates to stabilize 10.5 and the severity of the bugs found during that cycle and the resulting bad press. Leopard was a HUGE feature release and with it came tons of (bad) bugs.
The first .1 or .2 ALWAYS fixed critical bugs, because:
You had to GM / freeze the software to physically create the CDs/DVDs around a month before the release. Bugs found after this process required a repress (can’t remember the phrase we used), which cost money and time and scrambled effort at the last minute and added risk. This means the bar was super high, and most “bad, but not can’t use your computer bad” bugs were put in software updates…which was developed concurrently with the end of the main release (hence why .1 came out right away)
Testing was basically engineers, internal QA, some strategic partners like Adobe and MS, and the Apple Seed program (which was tiny). There was very little automated testing. Apple employees are not representative of the population and QA coverage is never very complete. And we sometimes held back features from seed releases when we were worried about leaks, so it wasn’t even the complete OS that was being tested.
Software updates are always needed, though the issues they fix became less severe over time due to larger seeds (aka betas), recovery partitions, and better / more modern development practices. But I can tell you FOR A FACT that Snow Leopard had fewer major bugs over its lifetime, coalesced very quickly, and was extremely solid when Lion was released.
LegNeato
哦,又是这个。我真该搞个网站来放这个……
我就是当年在苹果负责10.6安全更新(10.6.1+)的"直接责任人"(DRI)。我和更新项目办公室的团队每天审查每一个bug,判断它应该放进安全与稳定性更新,还是等到下一个大版本再修。说真的,每天早晨我们都会对所有的Mac OS X bug进行分组分类,不管是新提交的,还是内部提名让我们评估是否应该加入更新的。我负责打包和审核构建,调整增量更新与完整更新的比例。我还构建了一套系统,大幅自动化了软件更新的差异对比"列车"(自动母版制作)。
新版本的OS总是通过分支/列车方式并行开发的,发现修复后会反向移植到当前版本。它们并不是线性逐个开发的。所以,如果你拿最稳定、打磨完善、修复停滞的最后一个大版本,与全新推出的1.0大版本分支相比,新的大版本当然bug更多。每个y.0版本相对于之前的x.8版本都是如此。但如果你比较的是大版本之间,雪豹确实不一样。
雪豹的内部目标明确标榜为减少bug、提升质量。这是事实,不是营销。我不明白为什么网上有些人就是不信,但我当时就在那里。你想加入一个功能,必须获得领导层的明确批准,门槛非常高。在正常的特性版本中,运作方式是自下而上的"这是我们计划要发布的",而雪豹是自上而下的"我们可以发布这个吗?"
据我所知,雪豹是第一个这种类型的版本(我参与的第一个版本是Jaguar或Puma),它直接回应了10.5需要8个软件更新才能稳定下来、以及那个周期中发现bug的严重性和随之而来的负面报道。Leopard是一个超大的特性版本,随之而来的是大量(糟糕的)bug。
第一个.1或.2版本总是会修复关键bug,因为:
你必须在大约发布前一个月完成GM/冻结软件,以便物理生产CD/DVD。在此之后发现的bug需要重新压制(我不记得我们用的术语了),这既费钱又费时间,还会在最后一刻打乱工作节奏并增加风险。这意味着门槛非常高,大部分"糟糕但还不至于让电脑没法用"的bug被放进了软件更新……而软件更新是与主版本发布末期同时开发的(所以.1很快就能出来)。
测试基本上要靠工程师、内部QA、一些战略合作伙伴(比如Adobe和微软),还有苹果种子计划(规模很小)。自动化测试极少。苹果员工并不能代表全体用户,QA覆盖率也从来不是百分之百。而且有时候我们担心泄露,会在种子发布中推迟某些特性,所以测试的甚至都不是完整的操作系统。
软件更新总是需要的,不过随着种子(也就是测试版)规模扩大、恢复分区出现以及更好/更现代的开发实践,它们修复的问题严重程度随时间降低了。但我可以明确告诉你一个事实:雪豹在整个生命周期中重大bug更少,稳定得非常快,并且在Lion发布时已经极其稳固。
https://news.ycombinator.com/item?id=49155075
We can chalk this up as another example of over-exhuberance by what folks believe LLMs can accomplish vs. what they actually are.
LLM-based “AI” is able to use its vast corpus of inputs and calculate the most statistically likely output in a given situation. It is probabilistic, and when you are dealing with probabilities in a situation where certainties, not probabilities, matter, you’re going to get dinged on credibility massively when your LLM-based “AI” gets the probabilities wrong at best, or in this case, claims a line of code generates a vulnerability when it is, in fact, a code comment.
LLMs are text-prediction engines. They are not Artificial Intelligence, and shouldn’t not be treated in any form or fashion as if they possess intelligence. What bothers me about this entire situation is that presumably the folks that relied on the LLM-based “AI” to generate these vulnerabilities knew (or should have known) enough about their tool to know this would happen, but did not.
Now, we all pay the consequence, to the tune of hundreds of thousands if not millions of dollars of wasted productivity from teams that have to deal with the resulting fall-out of this usage of “AI”.
A human must verify everything an LLM presents as fact. Everything. If you don’t, we all pay the price. LLMs do not remove the onus of responsibility on the human being, if anything they amplify it because LLMs can generate lots more output more quickly that needs to be verified than humans can.
gortok
我们可以把这件事归为另一个例子,说明人们认为LLM能完成的事情与实际能力之间存在过度乐观的差距。
基于LLM的“AI”能够利用其庞大的输入语料库,计算在特定情境下最可能统计出现的输出。它是概率性的,而当你在处理需要确定性而非概率的情况时——就像这次,基于LLM的“AI”最多只是在概率上出了错,或者像本例中,它声称一行代码产生了漏洞,而实际上那只是一行代码注释——那么你的可信度就会遭受重创。
LLM是文本预测引擎,它们不是人工智能,绝不应以任何形式被当作拥有智能来对待。整件事让我感到不安的是,那些依赖基于LLM的“AI”来生成这些漏洞的人,本该(或应该)对自己的工具有足够了解,预见到这种情况会发生,但他们却没有。
现在,所有人都要为此付出代价——从那些不得不处理这种“AI”使用后果的团队所浪费的生产力来看,损失高达数十万甚至数百万美元。
人类必须核实LLM呈现为事实的一切信息。一切。如果你不这样做,我们所有人都会付出代价。LLM并没有减轻人类应承担的责任,反而放大了它,因为LLM能比人类更快地生成大量需要验证的输出。
https://news.ycombinator.com/item?id=49156719
One of the arguments for open source software for end-users has always been the freedom to examine and modify how that software works.
The reality for most people - even expert programmers - has been that the freedom is more about being able to lean on other people to do that. Most people can’t justify the time commitment needed to read and then modify the code for tools they use very often.
I think LLMs have changed that equation in a way that makes the original dream much more feasible.
Several times a day I’ll prompt regular Claude chat to “Clone x/y from GitHub and tell me how Z works”.
Getting software to compile in order to start hacking on it used to be enough friction that I often wouldn’t bother. Now I treat that as a zero time investment challenge: tell Codex or Claude Code to checkout and build X and then come back ten minutes later and see how it got on.
I’m not habitually modifying the software I use yet, but I can see a path to that which didn’t exist a year or so ago.
simonw
开源软件对终端用户的一个论据一直是:可以自由地检查并修改软件的工作方式。
但对大多数人——甚至包括专家级程序员——而言,这种自由实际上更多是能够依赖他人来完成这些工作。大多数人无法证明花时间阅读并修改他们经常使用的工具代码是值得的。
我认为大语言模型改变了这一局面,使得最初的理想变得可行得多。
每天我会好几次让普通的Claude对话“从GitHub克隆x/y并告诉我Z是如何工作的”。
要让软件能够编译以便开始修改代码,过去这往往是一个麻烦到让我经常懒得去做的步骤。现在我把这当作一个零时间投入的挑战:让Codex或Claude Code检出并构建X,然后十分钟后回来看看它进展如何。
我还没有养成习惯性地修改我所用软件的习惯,但我能看到一条通向这种习惯的路径——而这条路在一年前左右还不存在。
https://news.ycombinator.com/item?id=49153521
On social media, I saw the much more vulgar
“Learned engineering just to become the condom between Claude Code and prod”
And that (re)framing helped as well to think about the “what are we even (left) doing” as an industry
gregsadetsky
在社交媒体上,我看到更粗俗的说法:“学了工程学,结果只是成了Claude Code和生产环境之间的避孕套。”而那个(重新)框架也帮助思考“我们这个行业到底(还)在做什么”。
https://news.ycombinator.com/item?id=49155229
Big no for retyping llm generated code by hand.
But a big yes for still typing code by hand, and not leaving it to the llm. Except it has to be the code generated by your brain.
That is what will create new neurons and new connections, which is what will keep away the cognitive decline.
And the constraint of not having to use llms will enhance creativity.
Actually, the constraints llms add to your code are more in number than the former. llms code in only the specific ways they’ve been trained on. So you won’t ever come across of other ways.
Off the top of my head.. here’s RubyQuiz.com [0] which I came across when I was learning ruby more than a decade ago. Looking at the many user-submitted solutions (you have to download the zip file!) you’ll see completely different ways the problems were solved.
Sure, many won’t be deemed efficient or standard by today’s llm or rubocop checks, but looking at their code.. and retyping them and seeing them work.. was crucial in how I was able to think in Ruby for solving coding problems.
I did the same with Go too, with the “learn go with tests” guide [1].
[0] - http://rubyquiz.com/
[1] - https://quii.gitbook.io/learn-go-with-tests
npras1
强烈反对手动重新输入LLM生成的代码。
但强烈赞成仍然手动输入代码,而不是把它交给LLM。只不过这些代码必须是由你的大脑生成的。
这样才能创造出新的神经元和新的连接,从而延缓认知衰退。
而且不必使用LLM这一限制会增强创造力。
实际上,LLM给你的代码增加的限制比前者更多。LLM只会按照它们被训练过的特定方式编写代码。所以你永远无法接触到其他方法。
凭记忆想到……这里有个RubyQuiz.com[0],是我十多年前学Ruby时遇到的。看看那些用户提交的众多解决方案(你需要下载zip文件!),你会看到人们用完全不同的方式解决这些问题。
当然,很多方案在今天被LLM或Rubocop检查时可能不被认为是高效或标准的,但看看它们的代码……然后重新输入它们,看着它们运行……这些对于我能够用Ruby思维解决编程问题至关重要。
我对Go也做了同样的事情,用的是"learn go with tests"指南[1]。
[0] - http://rubyquiz.com/
[1] - https://quii.gitbook.io/learn-go-with-tests
https://news.ycombinator.com/item?id=49158545
The CDC Has a Cyclospora Lab. DOGE Downsized It Last Year [1]
[1] https://www.wired.com/story/cdc-cyclospora-lab-doge-downsized-it-last-year/
wnevets
美国疾控中心有一个环孢子虫实验室。去年DOGE将其规模缩减了[1]
https://news.ycombinator.com/item?id=49149976
In summary: Because OSM requires work and care to be put into the data submission plan, which isn’t worth it.
A project like OSM would be bombarded with spam and junk submissions if it didn’t have these barriers to submission. Understandable.
Aurornis
总结:因为OSM要求提交数据计划时投入大量精力和细心,而这并不值得。
像OSM这样的项目,如果没有这些提交门槛,就会被垃圾信息和无用提交所淹没。这可以理解。
https://news.ycombinator.com/item?id=49149190
Asking “what have note-taking apps accomplished?” is like asking “what have spreadsheets accomplished?” or “what have cameras accomplished?”. A tool doesn’t accomplish anything on its own. A tool is a means to an end.
A few days ago, someone jokingly asked if Henry Ford had a personal knowledge base in Obsidian… Well, sort of! He had something he called “jot books”, where he journaled, kept notes, grocery lists, etc. Not dissimilar to how people use Obsidian. The Henry Ford Museum has fifty of these notebooks: https://www.thehenryford.org/search?Query=%22jot+book%22
How should we quantify the impact of Henry Ford’s notebooks? How should we quantify the impact of spreadsheets?
Most people who accomplish anything take notes in some form, because writing is a way of thinking. We love to mythologize the tools and methods of accomplished people because we hope it will let us absorb a bit of their genius. But a good camera doesn’t make a good photographer. Taking lots of photos helps.
Should you take notes? Probably. Does it matter what your method is? Probably not. Whatever works for you. Obsidian (or any other form of notetaking) is successful if it disappears and lets you accomplish your work.
kepano
问“笔记应用完成了什么?”就像问“电子表格完成了什么?”或“相机完成了什么?”一样。工具本身不会完成任何事情,工具只是达成目的的手段。
几天前,有人开玩笑地问亨利·福特是否在Obsidian里建有个人知识库……嗯,差不多吧!他有种自己称为“便签本”的东西,用来写日记、记笔记、列购物清单等,和人们使用Obsidian的方式没什么不同。亨利·福特博物馆收藏了五十本这样的笔记本:https://www.thehenryford.org/search?Query=%22jot+book%22
我们该如何量化亨利·福特笔记本的影响?又该如何量化电子表格的影响?
大多数有所成就的人都会以某种形式做笔记,因为写作本身就是一种思考方式。我们热衷于将成功人士的工具和方法神话化,希望借此吸收他们的一丝天赋。但好相机不会自动造就优秀摄影师——多拍照片才有帮助。
你应该做笔记吗?大概是的。你的方法重要吗?大概不重要。适合你的就行。当Obsidian(或其他任何笔记形式)隐形消失、让你专注于完成任务时,它才算真正成功。
https://news.ycombinator.com/item?id=49148914
All this ostensibly to keep teenage boys from watching Pornhub (when parental controls already exist).
The real reason, of course, is to force people to connect strong real-life identifiers to online activity. Mobile first, then Windows. Then Linux is too weak to oppose on its own, and will adapt or die.
big85
这一切表面上是为了防止青少年男孩观看Pornhub(而家长控制功能早已存在)。当然,真正的目的是迫使人们将强有力的现实身份标识与在线活动关联起来。先从移动端开始,然后是Windows。至于Linux,它太弱小了,无法独自对抗,要么适应,要么消亡。
https://news.ycombinator.com/item?id=49154534
The problem with this kind of thing, is that it reduces the S/N (Signal-to-Noise) ratio, so weeding out the legit CVEs becomes a lot more difficult.
But, on the other hand, I do know that LLMs have been discovering a lot of legit CVEs, and I will lay odds that the blackhats are leveraging them to the max.
ChrisMarshallNY
这类问题在于它降低了信噪比(S/N),导致筛选出真实的CVE变得更加困难。
但另一方面,我确实知道大型语言模型已经发现了大量真实的CVE,而且我敢打赌,黑帽黑客们正在最大限度地利用它们。
https://news.ycombinator.com/item?id=49147851
Brian Gilbert, 56, of San Jose, Calif., former Senior Manager of Special Operations for eBay’s Global Security Team, was sentenced to time served, one year of supervised release with the special condition that he have no contact with either of the victims in the case and a $20,000 fine
Jim Baugh, 47, of San Jose, Calif., eBay’s former Senior Director of Safety and Security, was sentenced to 57 months in prison
David Harville, 50, of Las Vegas, Nev., former Director of Global Resiliency, was sentenced to 24 months in prison
Stephanie Popp, 34, of Louisville, Ky., former Senior Manager of Global Intelligence, was sentenced to 12 months in prison
Philip Cooke, 56, of San Jose, Calif., a former Senior Manager of Security Operations, was sentenced to 18 months in prison and 12 months of home confinement
Stephanie Stockwell, 28, of Redwood City, Calif., a former Manager of Global Intelligence, was sentenced to one year in home confinement
Veronica Zea, 28, of San Jose, Calif., a contract intelligence analyst, was sentenced to one year in home confinement
https://www.justice.gov/usao-ma/pr/final-defendant-ebay-cyberstalking-case-sentenced
haunter
56岁的布莱恩·吉尔伯特,来自加利福尼亚州圣何塞,曾担任eBay全球安全团队特别行动高级经理,被判处已服刑期、一年监督释放(附加条件为不得接触本案任何受害者)及2万美元罚款。
47岁的吉姆·鲍,来自加利福尼亚州圣何塞,eBay前安全与安保高级总监,被判处57个月监禁。
50岁的戴维·哈维尔,来自内华达州拉斯维加斯,前全球韧性总监,被判处24个月监禁。
34岁的斯蒂芬妮·波普,来自肯塔基州路易斯维尔,前全球情报高级经理,被判处12个月监禁。
56岁的菲利普·库克,来自加利福尼亚州圣何塞,前安全运营高级经理,被判处18个月监禁及12个月居家监禁。
28岁的斯蒂芬妮·斯托克韦尔,来自加利福尼亚州雷德伍德城,前全球情报经理,被判处一年居家监禁。
28岁的维罗妮卡·塞亚,来自加利福尼亚州圣何塞,合同情报分析师,被判处一年居家监禁。
https://www.justice.gov/usao-ma/pr/final-defendant-ebay-cyberstalking-case-sentenced
https://news.ycombinator.com/item?id=49154911
The vast majority of CVEs are not exploitable, basically noise. I suspect that the overwhelming majority of the CVEs being generated by LLMs are either noise of the sort in the linked article or noise of the sort that is not exploitable.
flerchin
绝大多数CVE(通用漏洞披露)并不可被利用,基本上就是噪音。我怀疑由LLM生成的海量CVE中,要么是链接文章中那种类型的噪音,要么就是不可被利用的那类噪音。
https://news.ycombinator.com/item?id=49146705
A lot of people are posting here about how bad the end product is, but that is kind of the point. Models have moved beyond generating images to a new kind of benchmark that better exposes understanding of the physical world, and we can use benchmarks like this to measure future progress. (Of course, it will have to be a qualitative/subjective measurement.)
jmugan
很多人在这里吐槽最终效果有多差,但这恰恰是重点所在。模型已经从生成图像迈向了新的评测基准,这种基准能更好地暴露对物理世界的理解程度,而我们也可以用像这样的基准来衡量未来的进步。(当然,这必然会是一种定性的/主观的评估方式。)
https://news.ycombinator.com/item?id=49153613
All requests to an LLM are idempotent, for every API call you need to send it the entire conversation history
A more appropriate term is “stateless”. LLM responses are certainly not idempotent, as they are not even deterministic.
jcoder
所有对LLM的请求都是幂等的,因为每次API调用都需要发送整个对话历史
更合适的术语是“无状态”。LLM的响应当然不是幂等的,因为它们甚至不是确定性的。
https://news.ycombinator.com/item?id=49148089
This article is a bit rambly so I’ll just focus on some things from the beginning:
If your kitchen knife kept changing shape, weight, and edge, you’d have to relearn it every time; that’s a hard tool to build trust in.
This concept was betrayed far before agentic tools, with a much earlier concept: Automatic updates.
To use one product as an example: When Windows ME and Windows Vista came out, people hated them even more than they usually hated Windows, so they did not use them. Microsoft was forced to respond by making a not-quite-as-bad OS in Windows XP and a pretty good OS in Windows 7 respectively. No longer is that an option, your workflow will simply be interrupted by automatic updates.
Vim and Emacs, in their infinite customizability, can be molded to fit your exact hand and workflow
Vim is one major exception to the automatic update problem. I trust vim not just because it can do a ton of shit (although that is certainly nice), but because unlike most other software, its UI doesn’t change unless I tell it to change. Aside from switching from vim to neovim (my decision, not a forced update), my muscle memory from a couple decades ago still works today.
MiddleEndian
这篇文章有些散乱,所以我只聚焦开头部分的内容:
“如果你的菜刀不断改变形状、重量和刀刃,你每次都得重新适应它——这样的工具很难让人信任。”
早在智能代理工具出现之前,这个概念就已经被一个更早的概念背叛了:自动更新。
以某个产品为例:当Windows ME和Windows Vista发布时,人们对它们的厌恶程度甚至超过了以往对Windows的惯常反感,因此没人愿意用它们。微软被迫做出回应,分别推出了还算不错的Windows XP和相当优秀的Windows 7。如今这种选择权已不复存在,你的工作流程随时会被自动更新打断。
“Vim和Emacs拥有无限的定制性,可以被塑造成完全贴合你的手型和工作流程。”
Vim是自动更新问题的一个重大例外。我信任Vim,不仅因为它能做海量的事情(虽然这确实很棒),更因为它与大多数软件不同,它的用户界面不会改变——除非我主动要求它改变。除了从Vim切换到Neovim(这是我自己的决定,而非强制更新),我几十年前形成的肌肉记忆至今依然有效。
https://news.ycombinator.com/item?id=49149937
Gutenberg’s copy of Moby Dick is 1.2MB[0]. Which is to say the slowest benchmarked terminal could display a paltry ~53 Moby Dicks per second, while shitty gives you ~98 Moby Dicks.
I am not sure how many Moby Dicks I require per second, but it is good to have options.
[0] https://www.gutenberg.org/ebooks/2701
3eb7988a1663
古腾堡版《白鲸》的文本大小为1.2MB[0]。也就是说,最慢的基准测试终端每秒也只能显示区区约53本《白鲸》,而最烂的终端每秒能显示约98本。
我不确定自己每秒需要多少本《白鲸》,但有选择总归是好事。
[0] https://www.gutenberg.org/ebooks/2701
https://news.ycombinator.com/item?id=49147991
No consequences for executives, and yet their ostensibly enormous responsibility is how their salaries and benefits are always justified.
happytoexplain
高管们无需承担任何后果,然而他们表面上巨大的责任却总是被用来证明其薪资和福利的合理性。
https://news.ycombinator.com/item?id=49152488
At my last job, a coworker did this to me. The first time it happened, I ignored it. The second time, I responded in public saying “thanks but I can ask Claude myself.” Nobody ever pasted me an LLM response again. YMMV with team size and seniority though
maccard
在我上一份工作中,有位同事对我做了这样的事。第一次发生时我没理睬。第二次我公开回应说:“谢谢,但我可以自己去问Claude。” 之后再也没有人给我粘贴过LLM的回复了。不过具体情况可能因团队规模和资历而异。
https://news.ycombinator.com/item?id=49150929
Qwen3.6-35B is my daily driver for AI, and what convinced me to cancel my Claude subscription back in April. The Qwen3.6 line is easily the best local model I’ve tried, and I’ve tried a lot. I’ve got it diligently grinding away on my laptop right now, reviewing and fixing some bugs in my F# code.
nozzlegear
Qwen3.6-35B是我日常使用的AI模型,也是让我在四月份取消Claude订阅的原因。Qwen3.6系列绝对是我试用过的最佳本地模型——我试过很多。现在它正在我的笔记本电脑上勤恳工作,审查并修复我F#代码中的一些漏洞。
https://news.ycombinator.com/item?id=49146032
I tried to organize vocab by difficulty level for an English language-learning app once.
It shocked me how there is absolutely no “right” answer.
If you are teaching English for travel, then you’re prioritizing a lot of stuff around bathrooms, transportation, menu items, etc.
If it’s for understanding TV, it’s a lot of words like “murder”, etc. Depending on which TV shows you want to understand.
If it’s for reading the newspaper, you don’t ever need to know “bathroom”, but you sure do need to know words like “congressman”.
While if you are living somewhere, it’s really important to know a lot of basic supermarket items that you wouldn’t prioritize for other usages.
Also, while it’s easy to calculate word frequencies for stuff like newspaper articles, there aren’t any good statistics (last I checked) around just normal everyday conversation. Because that stuff isn’t getting recorded and transcribed. And the substitutes – transcribed speech from TV, radio, podcasts, etc. – is not the same context as the random stuff you say at home and during an average day.
crazygringo
我曾尝试为一个英语学习App按难度级别整理词汇。
但让我震惊的是,根本不存在所谓“正确”的答案。
如果你教的是旅行英语,那你会优先考虑关于卫生间、交通、菜单等词汇。
如果是为了看懂电视节目,那就会有很多像“谋杀”之类的词——具体取决于你想看懂哪些剧。
如果是读报纸,你根本不需要知道“卫生间”,但必须知道“国会议员”这类词。
而如果你住在一个地方,了解大量超市常见物品就非常重要,这些词在其他用途中却不会优先考虑。
此外,虽然计算报纸文章等文本的词频很容易,但关于日常普通对话的可靠统计数据(据我所知)却几乎没有。因为这类对话不会被录下来转写成文字。而替代品——电视、广播、播客等转写文本——和你在家、日常闲聊的语境完全不同。
https://news.ycombinator.com/item?id=49156731
I agree that devtools should be open source, but… I very much disagree with the premise that no tools should have config files, options, or plugin systems, and instead when you want to change something like your text editor’s font size, you should have an LLM download the code, change the hard-coded value, and rebuild it.
That’s just so inefficient and wasteful. Assuming a world in which LLMs do most of the coding work, do we want to burn electricity having the LLM build an options dialog or config file parser once, or do we want to burn electricity millions of times as users want to change any little thing about the software they use?
I really hope what you should expect is my answer to that question isn’t controversial.
Having an LLM do bespoke customizations that are unlikely to be interesting to other people? Great, sure. But adding a generally-useful feature to a piece of software, but not caring to try to upstream it? Lame. Lame, lame, lame.
kelnos
我同意开发者工具应该是开源的,但我非常不同意"任何工具都不该有配置文件、选项或插件系统,而当你想要改变像文本编辑器字号这样的东西时,应该让LLM下载代码、修改硬编码值并重新构建"这一前提。
这实在是太低效和浪费了。假设在一个LLM承担大部分编码工作的世界里,我们是想要烧一次电让LLM构建一个选项对话框或配置文件解析器,还是想要在用户每次想改变他们所用软件的任意小细节时,烧几百万次电?
我真的希望你能预期到,我对这个问题的回答并不会引发争议。
让LLM做那些不太可能对其他人有意义的定制化修改?很好,当然可以。但给一款软件添加一个普遍有用的功能,却不尝试将其向上游提交?那太差劲了。差劲,差劲,真差劲。
https://news.ycombinator.com/item?id=49151968
If you create a machine for laziness you’re going to get lazy people. It’s only going to get worse I’m afraid.
Do you guys think we’re going to see a de-evolution of human beings due to technology?
jpnc
如果你为懒惰创造机器,就会得到懒惰的人。恐怕情况只会越来越糟。你们认为我们会因为技术而看到人类退化吗?
https://news.ycombinator.com/item?id=49148504
I don’t understand where the all the EU anti-trust and anti-corruption regulators are here. Governments enforcing that you have a Google or Apple account to participate in society is transparently absurd.
This isn’t only a digital sovereignty issue, it’s also an anti-competition issue.
afandian
我不明白欧盟的反垄断和反腐败监管机构都在哪里。政府强制要求你拥有谷歌或苹果账户才能参与社会活动,这明显是荒谬的。这不仅是一个数字主权问题,也是一个反竞争问题。
2026-08-03 07:54:23
- 谷歌通过“拥抱、扩展、消灭”策略在多个产品中反复引入又移除RSS功能,严重损害了开放互联网的普及与信任。
- Diátaxis 方法将技术文档分为教程、操作指南、技术参考和解释四种形式,围绕用户需求提供系统化指导,有助于写出清晰且易维护的文档。
- 字节跳动发布Seedance 2.5视频生成模型,支持单次30秒生成、多模态引用和高级编辑功能,已上线多个平台,API即将推出。
- Karpathy使用Opus 5模型生成5500行代码渲染《指环王》场景,展示了LLM定制虚拟世界的潜力,但也暴露了模型缺乏多模态感知和实时交互的局限。
- Go 1.27引入泛型方法、结构体字面量字段选择、更快内存分配等新特性,本文通过可运行示例帮助开发者快速掌握变化。
- MIT研究发现AI财务建议整体质量不错,但过度依赖经验法则且提示词会因用户性别和素养不同导致建议差异,可能加剧贫富差距。
- 维基媒体基金会拒绝承认员工工会并聘请破坏工会的律所,引发社群不满和请愿,同时推出了限制候选人资格的董事会选举规则。
- 一位15岁爱好者展示了自行设计并3D打印的摆线齿轮箱,通过三个版本迭代实现1:9减速比,并分享了完整的工程实践。
- NetBSD 11.0正式发布,提供安装镜像并因AI导致安全问题增多而公开已知漏洞,修复计划在11.1版本推出。
- 科技记者批评Google News搜索功能因AI分心而严重退化,认为谷歌已放弃维护该产品,导致搜索结果充斥无关内容。
https://openrss.org/blog/how-google-helped-destroy-adoption-of-rss-feeds
Google 在多个产品中反复引入 RSS 功能后又将其移除或限制,严重影响了 RSS 的普及和用户信任。其行为包括:在 Chrome 浏览器中悄然取消内置 RSS 按钮;收购 FeedBurner 后关闭 API 并大幅削减服务;关闭广受欢迎的 Google Reader,导致大量用户放弃 RSS;从 Google Alerts 中短暂移除 RSS 后又恢复;误删官方 RSS 扩展后又重新上架;以及彻底停止 Google News 的 RSS 支持。文章指出,Google 利用开源 RSS 协议吸引用户,再逐步锁定并抛弃,这种“拥抱、扩展、消灭”的模式对开放互联网构成威胁。虽然 Google 在 2021 年宣布可能重启 Chrome 的 RSS 功能,但至今未正式推出,用户对其长期可靠性仍存疑虑。
https://news.ycombinator.com/item?id=49136821
这是一个介绍 Diátaxis(一种技术文档编写方法)的网站。Diátaxis 源自希腊语,意为“跨与排列”,主张从系统化角度理解文档用户的需求,并据此确定四种不同的文档形式:教程(tutorials)、操作指南(how-to guides)、技术参考(reference)和解释(explanation)。这四种形式被置于系统性的关系中,文档本身也应围绕这些需求结构来组织。
Diátaxis 能帮助解决文档的内容(写什么)、风格(怎么写)和架构(如何组织)问题。它不仅服务文档用户,对文档创建者和维护者也有价值:轻量、易掌握、易应用,不限制具体实现,并为文档质量提供主动性原则。
网站内容包括快速入门指南(Start here)、应用方法(如教程、指南、参考、解释的说明,以及“指南针”和工作流程)、深入的理论与原则(包括基础、地图、质量、复杂层级等话题)。该框架已在许多文档项目中得到验证,并获得了来自 Vonage、Gatsby 和 Cloudflare 等公司的实践者推荐,他们认为 Diátaxis 帮助团队明确了信息架构,使文档对读者和贡献者都更加清晰。
https://news.ycombinator.com/item?id=49138188
https://seed.bytedance.com/en/blog/one-take-creation-flexible-referencing-introducing-seedance-2-5
Seedance 2.5 是字节跳动发布的新一代视频生成模型,定位为“一次创作、灵活引用”。它基于 Seedance 2.0 的统一多模态音视频联合生成架构,围绕基础生成与引用生成进行升级,在长视频叙事、多模态引用和编辑能力上实现重要突破,致力于让用户从生成简单片段升级为完成完整创意作品。
主要亮点包括:单次生成时长从 15 秒提升到最长 30 秒,并支持多轮扩展,能够保持人物、场景、叙事节奏和视听风格的一致,方便用户连续拼接生成数分钟的高质量长视频;多模态引用能力全面增强,一次可输入最多 30 张图片、10 段视频和 10 段音频,并强化了粘土渲染、动态引用、创意引用等能力;编辑功能更精确稳定,支持时间戳级音视频定位修改,还增强了绿幕、机位、引用编辑等高级功能,适用于影视和广告等专业领域。
模型同时改善了 AI 视频中常见的“人工感”,在物体纹理、皮肤与眼睛细节、光照、色彩饱和度等方面进行优化,并减少字幕和背景音乐的不可控出现,使成片更接近真实电影感。官方还展示了由 Seedance 2.5 端到端制作的创意短片。
目前 Seedance 2.5 已在即梦 AI、豆包 Pro 等平台陆续上线,API 即将通过 BytePlus ModelArk 提供。项目主页为 https://seed.bytedance.com/seedance2_5。
https://news.ycombinator.com/item?id=49138302
https://twitter.com/karpathy/status/2083749667410727319
Andrej Karpathy 发布了一篇关于大型语言模型(LLM)测试新方向的帖子。他提到,人们已不再满足于“画一只骑自行车的鹈鹕”这类简单 SVG 测试。他尝试给 Opus 5 提供《指环王》开篇第一段,配上约 100 万 token(约 10 美元)的预算,要求用 three.js 渲染这段故事。Opus 5 花了约两个小时,生成了 5500 行代码,以程序化方式渲染出整个场景。
Karpathy 对此感到惊叹,因为 LLM 需要在三维坐标中放置并编排各种多边形资源,还要编写动画代码,而它居然真的做到了。他强调,这种高度定制化的内容以往根本没人愿意花时间制作,但 LLM 拥有无尽的耐心和精力,于是这类事情从“没人会做”变成了“几乎免费”的尝试。他期待未来能按需生成极度个性化的虚拟世界,例如让玩家进入《指环王》故事中,作为旁观 NPC 或某位角色参与其中,就像“按需生成的 GTA”一样。
不过他也指出了这类领域当前暴露出的局限:LLM 难以高效地原生观看视频或游玩游戏来审查自己的作品。在这个实验中,Opus 5 只能缓慢地截取不同节点的截图,还多次出错,产生了不少粗糙之处。他认为多模态和游戏内感知能力仍然非常欠缺。
最后,Karpathy 分享了他上传的源代码,可在浏览器中游玩和改编,地址为 karpathy.ai/lotr-movie/,并调侃说“GTA 夏尔”可能会比 GTA VI 更早问世。埃隆·马斯克也在下方回复了一个“Yah”表示赞同。
https://news.ycombinator.com/item?id=49140998
https://victoriametrics.com/blog/go-1-27/index.html
Go 1.27 即将发布,本文以可运行示例介绍其重要新特性。
主要更新包括:
GOEXPERIMENT=nosizespecializedmalloc 关闭。GODEBUG=tracebacklabels=0 禁用。goroutineleak 配置文件成为正式功能,用于检测永久阻塞的 goroutine。文中还提供了相关文档、提案、提交和作者链接,并推荐了此前 Go 1.22 至 1.26 的系列教程。
https://news.ycombinator.com/item?id=49140218
这是一篇来自 MIT 斯隆管理学院的报道,介绍了一项关于人工智能财务建议质量的新研究。
研究发现,AI(大型语言模型)提供的财务建议总体质量不错,能鼓励人们增加储蓄、投资多样化,并随着年龄增长降低股票风险。尤其是 30 岁以上人群,遵循 AI 建议可以积累可观的储蓄缓冲。
但 AI 建议也存在明显不足:它过于依赖简单经验法则,在面对失业等冲击时调整不当,比如建议失业者过度削减开支;同时,AI 往往让投资组合自然漂移,而非主动再平衡。
研究还发现,提示词的质量很关键。与普通用户提问相比,结构更完整的“学术式提示词”能显著提升 AI 建议的质量。此外,AI 的回答会因提问者的性别、金融素养和 AI 使用经验而产生差异,遵循男性、金融素养较高或有 AI 使用经验者的提示,可让退休时财富增加约 5%,这可能引发贫富差距。
https://news.ycombinator.com/item?id=49139102
https://en.wikipedia.org/wiki/Wikipedia:Wikipedia_Signpost/2026-08-02/News_and_notes
维基媒体基金会拒绝了员工工会“美国维基工人联合会”的自愿承认请求,并聘请了以“破坏工会”闻名的律师事务所 Littler Mendelson。工会指责基金会的做法是拖延战术,并已向国家劳工关系委员会申请选举。社群对此普遍不满,已发起多份请愿支持工会。此外,基金会还提出了严格限制候选人资格的董事会选举规则,引发争议。
https://news.ycombinator.com/item?id=49143414
https://github.com/tom-ilan/cycloidal_gearbox
这个网页是一个关于摆线齿轮箱(Cycloidal Gearbox)的 GitHub 仓库,作者展示了自行设计和 3D 打印制作该齿轮箱的过程,并提供了用于生成齿轮参数的 Python 脚本。
设计过程分为三个版本:
Python 脚本基于 SolidWorks 相关文章中的摆线方程编写,可生成摆线针轮轮廓,支持调节针销数、偏心距、针销半径等参数,并可加入 3D 打印间隙。
版本 3 的关键规格:齿轮比 1:9(10 根外针销,9 个转子凸角),外径 90 mm,使用 NEMA 17 步进电机,PLA 材料打印,配备 4 个 M3×8 螺丝和 2 个 6704 轴承,公差偏移 +0.15 mm。输出扭矩约 1.3 N·m,电机基础扭矩约 0.21 N·m,效率约 66%。
后续改进方向:将外壳针销替换为 MR128 轴承降低摩擦,输出针销改为带金属套的 M2 螺丝以提高刚性和扭矩。
https://news.ycombinator.com/item?id=49140396
https://blog.netbsd.org/tnf/entry/netbsd_11_0_released
NetBSD 11.0 正式发布!该版本历经较长时间开发,已提供安装说明和下载镜像。针对 ARM 设备有预配置 U-Boot 的镜像;ISO 已拆分为 CD 用的 700MB 镜像和完整 DVD 镜像,建议无尺寸限制时选择带“-dvd.iso”的版本。使用 U 盘等闪存介质需用 .img 文件并先解压。
关于安全漏洞:由于 AI 工具导致安全问题激增,项目组选择公开未解决的问题,不再推迟发布。当前已知问题包括:hdaudio(4) ioctl 缺少权限检查(可手动删除 /dev/hdaudio* 缓解)、ipfilter 空指针解引用(默认内核未启用)、pf 分片重组释放后使用(默认未启用)。这些修复将在 11.1 中提供,预计两个月内发布。
https://news.ycombinator.com/item?id=49136736
https://elgan.com/google-news-is-just-forrest-gumps-shrimp-boat-now
这是一篇科技记者 Mike Elgan 的博客文章,抱怨 Google News 搜索功能严重退化。他比喻说,Google 就像《阿甘正传》里被 Lieutenant Dan 分心而让船撞上码头的 Forrest,而 AI 就是那个“Lieutenant Dan”。
作者长期用 Google News 搜索美国新闻机构在特定时间段(如过去一周或一天)发布的报道,但如今这些设置逐渐失灵:搜索结果大量来自 Instagram 等社交平台而非新闻网站;很多结果是他无法识别的外语内容;许多文章来自国外而非他指定的美国;更离谱的是,选择“过去一周”后仍会给出数月甚至数年前的旧闻。
他认为 Google 已经放弃了对 Google News 的维护,任由其“撞毁”。文章以个人经历表达了对该产品衰退的失望。
https://news.ycombinator.com/item?id=49137681
https://news.ycombinator.com/item?id=49140187
People in this thread are massively underestimating the level of financial illiteracy in the general population.
We’ve had multiple people try to convince us to set up bank accounts for our kids, so that they could accumulate interest over 18 years.
More that tried to convince me to gamble on random pump and dump shitcoins.
More still that talked about “investing” in random collectables like Funko Pops or Pokemon cards - they’re not a bubble, Logan Paul told me so!
You could replace the AI with a piece of paper that says “set aside 10% of your income and invest it in an ETF” and it would outperform the financial “advice” that people receive on a daily basis.
AussieWog93
这个帖子里的很多人严重低估了普通人群体的金融知识匮乏程度。
我们已经遇到好几个人试图说服我们给孩子们开银行账户,这样他们就能在18年里累积利息。
还有更多人试图说服我去炒那些随机拉盘砸盘的垃圾币。
还有更多的人谈论“投资”像Funko Pop或神奇宝贝卡这样的随机收藏品——他们不是说这是泡沫,因为Logan Paul告诉我不是!
你甚至可以用一张写着“把你10%的收入存起来,投资到ETF里”的纸来取代AI,它也会比人们每天收到的那些金融“建议”表现更好。
https://news.ycombinator.com/item?id=49138695
The Internet in the early 2000s felt a little more special as compared to today, when 99.999 % of content is locked in a handful of walled gardens. I mean sure everything we had then is still possible but let’s be honest everything about the web is fine tuned to deliver ads to our eyeballs, browsers, operating systems and even protocols are being designed with ad delivery in mind. Even the privacy friendly browsers just exist to serve ads when I think about it. Kind of sad really, but I guess you could be happy as it’s so big now and there’s so much to do. Just makes you wonder what stuff would thrive if ads didn’t exist. As of now every niche thing that gets successful will be invariably pulled into the ad ecosystem as it’s very hard to say no to money.
RSS wasn’t in the interest of the big platforms as it’s decentralized and there’s no good way to deliver ads through it, simple as that.
throwaway63467
与今天相比,21世纪初的互联网感觉更特别一些,当时99.999%的内容都被锁在少数几个封闭的花园里。我的意思是,当然我们当时拥有的一切如今仍然可能实现,但说实话,现在的网络一切都是为了把广告送到我们眼前而精心调校的,浏览器、操作系统,甚至协议都在以广告投放为设计目标。仔细想想,就连那些注重隐私的浏览器,存在意义也只是为了投放广告。想想还真有点可悲,不过我想你也可以感到高兴,因为网络现在这么大,有那么多事情可做。只是让人不禁想,如果广告不存在,什么东西会蓬勃发展。而如今,任何小众的东西一旦成功,都必然会被拉进广告生态系统,因为很难对金钱说不。
RSS不符合大平台的利益,因为它是去中心化的,而且没有很好的方式通过它投放广告,就这么简单。
https://news.ycombinator.com/item?id=49140558
(I was on the Go team for ages)
Seriously, that’s all it was. Just Ian alone proposed and rejected a half dozen of his own different approaches to generics. Finally a language + implementation plan came together that people all liked.
Nobody was ever opposed to generics that I saw.
bradfitz
(我在Go团队待了很久) 说真的,事情就是这样。仅伊恩一人就提出并否决了自己六种不同的泛型方案。最终,一个大家都喜欢的语言+实现方案成形了。据我所见,从来没有人反对过泛型。
https://news.ycombinator.com/item?id=49147636
I don’t think it’s a bad way to benchmark new models, I just find it concerning that the author implies that “pelican on a bicycle” has been exhausted. At the risk of making overly broad, unfalsifiable claims I think multi-year exposure to AI content has dramatically raised our expectations for speed and volume but lowered them for quality. We see a very janky pelican and declare the problem solved.
YmiYugy
我不认为这是给新模型做基准测试的坏方法,我只是觉得令人担忧的是,作者暗示“骑自行车的鹈鹕”已经被用尽了。冒着提出过于宽泛、无法证伪的主张的风险,我认为多年接触AI内容已极大地提高了我们对速度和数量的期望,却降低了对质量的期望。我们看到一只非常蹩脚的鹈鹕,就宣布问题解决了。
https://news.ycombinator.com/item?id=49138079
Google’s obviously fake excuse for killing their RSS reader (declining usage) was especially maddening at the time because they were pushing Google+ - which nobody used.
bakemawaytoys
谷歌关闭其RSS阅读器时给出的借口(使用量下降)显然是假的,当时尤其令人恼火,因为他们正在全力推广谷歌+——而那个服务根本没人用。
https://news.ycombinator.com/item?id=49132926
In a way the most remarkable thing about this is that it isn’t even at the top of the HN homepage. Even if this is a step up from what we’ve seen before, we’re no longer astonished by the idea that AI can make significant advances in mathematics and computer science.
robinhouston
从某种意义上说,这件事最了不起的地方在于,它甚至都没登上HN首页头条。即便这相比我们以前见过的有所进步,我们也已经不再对AI能在数学和计算机科学领域取得重大进展这一想法感到震惊了。
https://news.ycombinator.com/item?id=49137277
I know that we’re discouraged from meta-comments, but what is going on in this thread? It’s a nearly 800-page book about the art of programming. A huge amount of work on a topic that should be dear to our hearts. News for hackers, right?
But somehow, the discussion has three themes. It’s 50+ comments of “I don’t like the first sentence of the marketing copy”, “I don’t like the tool the author is using”, and “what would happen if we train an LLM on this book?”. Has anyone read the sample chapter? Did you like it? Anyone here owns volume 1 and has opinions about that?
skippyfish
我知道我们被鼓励不要发元评论,但这个帖子里到底是怎么回事?这是一本将近800页的、关于编程艺术的书。在一个本应让我们牵肠挂肚的主题上投入了巨量工作。对黑客来说是个新闻,对吧?
但不知怎的,讨论却围绕三个主题展开。五十多条评论都是“我不喜欢营销文案的第一句话”、“我不喜欢作者用的工具”和“如果我们拿这本书训练一个大语言模型会怎样?”。有人读过样章吗?你喜欢吗?这里有人拥有第一卷,并且对它有看法吗?
https://news.ycombinator.com/item?id=49141859
This: “(b Box[T]) Map[U any](f func(T) U) Box[U]” is the type of cognitive weight I was happy that Go avoided.
baalimago
这个:“(b Box[T]) Map[U any](f func(T) U) Box[U]”是我很高兴 Go 语言避免的那种认知负担。
https://news.ycombinator.com/item?id=49142335
This is such a bikeshedding debate. While you don’t recommend it, projects with Tailwind work. Over years. You can onboard new developers to it, able to contribute productively immediately. Likewise, you can pick up work after months or years and don’t have to remember or rediscover how your styling layer works.
The conventions and class names come really naturally fast, and you can always look it up. It’s just not as a big of a problem people make it.
But the most ridiculous part of the article I found the cascade complaint:
<p class=“text-red-500 text-green-500”>I am some text</p>
Yes, this does not work. Why should it?! There is not a single use case where this is a good idea! In classic CSS, you might want to override something based on modifier classes, but that is just not a thing with Tailwind! If you end up programmatically layering class names, you’re looking at a code smell. Instead, you want to use attribute or state modifiers, like aria-hidden:opacity-0.
9dev
这是一场纯粹的“自行车棚”式争论。虽然你不推荐,但使用 Tailwind 的项目确实能跑起来,而且能跑很多年。你可以让新开发者快速上手,立刻就能高效产出。同样,隔了几个月甚至几年再捡起某个项目,你也不用去回忆或重新琢磨你的样式层是怎么组织的。
那些约定和类名学起来非常自然、很快就能掌握,而且随时可以查。这真的不像人们说的那么严重。
但我觉得文章里最离谱的是关于层叠(cascade)的抱怨:
<p class="text-red-500 text-green-500">I am some text</p>
是的,这样确实不会生效。但这本来就该这样啊?!根本没有任何一个场景会让人觉得这么写是个好主意!在传统 CSS 里,你可能想根据修饰类(modifier class)来覆盖某些样式,但 Tailwind 里压根没有这种用法!如果你最后要编程式地拼接类名,那本身就是代码坏味道了。你应该用属性或状态修饰符,比如 aria-hidden:opacity-0。
https://news.ycombinator.com/item?id=49132251
What remains to be seen is whether Google also introduced more Chrome bugs in June than over the past two years, thanks to AI.
The big problem is that AI output can be very convincing and look “right”, even appear to work, until you examine it in detail and realise all the edge-cases it didn’t handle.
userbinator
有待观察的是,谷歌是否也在6月借助AI引入了比过去两年更多的Chrome漏洞。
最大的问题在于,AI的输出可能非常有说服力,看起来“正确”,甚至看似能用,直到你仔细检查后才发现它没有处理所有边缘情况。
https://news.ycombinator.com/item?id=49142474
I see you have a .button, cool! So did you load the entire context of your project into your mind, and calculate every possible iteration of kind, size, color etc this button may have? And once you did that, did you come up with a semantically correct naming scheme that is clear and will not succumb to the inevitable .button_checkout_special_page_cta_widget a particular page will end up requiring?
No? Neither did I. I stopped thinking about CSS entirely almost a decade ago. Thanks Tailwind.
pixard
我看到你有一个.button,不错!那么你是不是把整个项目的上下文都装进脑子里,计算了这个按钮可能拥有的每一种类型、尺寸、颜色等等的所有迭代?而且一旦你算完了,你是不是想出了一个语义正确、清晰明了的命名方案,而且不会屈服于某个特定页面最终必然会需要的那个.button_checkout_special_page_cta_widget?
没有吧?我也没想过。我差不多十年前就不再纠结CSS了。感谢Tailwind。
https://news.ycombinator.com/item?id=49144574
Some people here think Wikimedia Foundation’s mission is to keep Wikipedia up. That’s a subset of their mission. I attribute nearly all of the blame to WMF, because their donation ads are quite deceptive.
The actual, public mission of Wikimedia Foundation:
The mission of the Wikimedia Foundation is to empower and engage people around the world to collect and develop educational content under a free license or in the public domain, and to disseminate it effectively and globally.
As such, WMF spends a lot money (combined) on a splatter of projects, like funding photographers to go to events like Fifa World Cup and Cannes, and take (CC or public domain) portraits for Wikimedia Commons ( https://www.wikiportraits.org ); etc.
dannyw
这里有些人认为维基媒体基金会的使命是维持维基百科运行。那只是他们使命的一部分。我几乎把所有责任都归咎于维基媒体基金会,因为他们的捐款广告相当具有欺骗性。
维基媒体基金会实际公开的使命:
维基媒体基金会的使命是赋权和吸引世界各地的人们,在自由许可或公有领域下收集和发展教育内容,并将其有效地在全球传播。
因此,维基媒体基金会将大量资金(合计)花费在各类项目上,比如资助摄影师前往国际足联世界杯和戛纳等活动,为维基共享资源拍摄(CC或公有领域)肖像( https://www.wikiportraits.org )等等。
https://news.ycombinator.com/item?id=49139882
There is a woman on twitter who makes seedance videos of her and Dario from anthropic, they’re kinda weird but generally pretty high quality: https://x.com/CuiMao/status/2058458683781365873 (full collection: https://x.com/CuiMao/status/2082740754380984373 ) - seeing them was the first time I’d been impressed with AI video gen.
neom
推特上有个女人用Seedance制作她和Anthropic的Dario的视频,虽然有点怪,但整体质量相当高:https://x.com/CuiMao/status/2058458683781365873 (完整合集:https://x.com/CuiMao/status/2082740754380984373)——看到它们是我第一次对AI视频生成感到惊艳。
https://news.ycombinator.com/item?id=49138514
Discs don’t matter. The whole “bring back discs” thing is pointless. They’re gone. Period.
Rights matter. If we had the same rights with digital purchases we had with old physical games the disc thing would be a much much smaller issue. Would many care outside true collectors?
Don’t confuse the two. If you do, and you complain loudly enough, you’ll get the monkey’s paw version of discs. All the downsides of both, no upside at all.
MBCook
光盘不重要。那些“把光盘带回来”的呼吁毫无意义。它们已经消失了。句号。
权利才重要。如果我们在数字购买上拥有和过去实体游戏一样的权利,光盘问题就会小得多。除了真正的收藏家,还有多少人在意呢?
不要把这两者混为一谈。如果你混为一谈,并且叫得够响,你得到的只会是“猴爪”版的光盘——集合了两者的所有缺点,没有任何好处。
https://news.ycombinator.com/item?id=49137979
Google Reader going away felt like the beginning of the end the internet as I knew it. I miss websites
betenoire
Google Reader的消失让我感觉像是我所熟知的那个互联网开始走向终结。我想念那些网站。
https://news.ycombinator.com/item?id=49132235
My main gripe here is the lack of transparency around the total experiment and construction. I doubt that they simply pointed their model at these ten specific problems alone and gave the model one shot; therefore the $2000 number could be completely misleading, similar to P-value hacking by not disclosing the total experimental setup.
I want to know:
aabhay
我主要的不满在于整个实验和构建过程缺乏透明度。我怀疑他们并不是仅仅把模型对准这十个特定问题,然后只给模型一次机会;因此,2000美元这个数字可能具有完全的误导性,类似于没有公开整个实验设置而进行P值操纵。
我想知道:
https://news.ycombinator.com/item?id=49131180
Even in the Era AI, GGPlot’s API is still the best charting API. The name “Grammar of Graphics” isn’t just marketing, they literally sought to write a god damn grammar to was capable of expressing all possible qualitative graphics.
They even wrote a book about how they went about it (not that it speaks to the quality of the API) https://link.springer.com/book/10.1007/0-387-28695-0
I actually stumbled upon this book when I was trying to look up how draftsmen (with pens and pencils on paper) did qualitative graphics as I found they had a lot of charm as opposed to modern charting libraries. It’s something I noticed when looking through a bunch of historical RBA (Reserve bank of Australia) annual reports, the 1960-1980 charts had a lot of character, but then you go into the early 2000s and its a stale chart from excel.
Anyways ggplot doesn’t really recapture the magic of those older charts, but it seems use quite a few of those as a baseline for how to communicate information. Like in figure 20.1 they talk about efforts to replicate older inforgraphics that showed Napoleon’s March on Russia, this graphic here (I think the example in the book is a bit nicer than the one in this blogpost IMO)
https://www.andrewheiss.com/blog/2017/08/10/exploring-minards-1812-plot-with-ggplot2/
On top of the charts just look nicer than anything you could produce with pyplot (and any API built on top of it) as pyplot seems to be have some really limited raster based rendering or something and the text handling is incredibly limited, I’ve never had this issue in ggplot.
I feel like most software engineers aren’t exposed to because it exists in the R ecosystem which is more so data scientist, econometricians, statisticians and other quantitative data professions, but it definitely one of the nicer APIs and I wish more people in the node and python ecosystem copied their homework. I see vega’s full name is something to do with grammars, but idk it’s for the same reason.
akst
即使在AI时代,GGPlot的API仍然是最好的绘图API。“图形语法”这个名字不只是营销,他们真的是想写出一套该死的语法,能够表达所有可能的定性图形。
他们甚至还写了一本书来讲他们是怎么做到的(虽然这本书并不能说明API的质量)https://link.springer.com/book/10.1007/0-387-28695-0
我其实是偶然发现这本书的,当时我在查制图员(用笔和纸的那种)是怎么做定性图形的,因为我发现他们的作品相比现代绘图库有一种特别的魅力。这是我翻阅一堆澳大利亚储备银行(RBA)历史年报时注意到的,1960到1980年代的图表很有个性,但到了2000年代初,就变成了Excel里那种毫无生气的图表。
总之,ggplot并没有真正重现那些老图表的魔力,但它似乎确实用了不少老图表作为信息传达的基准。比如在20.1图里,他们讨论了如何复刻展示拿破仑远征俄罗斯的旧信息图,就是这张(我觉得书里的例子比这篇博客里的稍微好看一点):
https://www.andrewheiss.com/blog/2017/08/10/exploring-minards-1812-plot-with-ggplot2/
而且这些图表看起来就比你能用pyplot(以及任何基于它构建的API)做出来的好看得多,因为pyplot好像是基于某种非常受限的栅格渲染,文本处理也极其有限——我在ggplot里从来没遇到过这些问题。
我觉得大多数软件工程师都没接触过它,因为它存在于R生态里,而R更偏向数据科学家、计量经济学家、统计学家以及其他量化数据从业者。但它绝对是最优秀的API之一,我希望node和python生态里能有更多人抄它的作业。我知道vega的全名跟“语法”有关,但我不确定是不是同样的原因。
https://news.ycombinator.com/item?id=49144375
I’m hardly a fan of the WMF, but the headline is clickbait. The WMF has used a law firm called Jones Day for brand and trademark management for over a decade. The firm is one of the largest legal firms in the US and it also does union busting, but a) the WMF does not appear to have engaged them for that, and b) the relationship long predates the current kerfuffle.
decimalenough
我算不上维基媒体基金会的粉丝,但这个标题纯属标题党。维基媒体基金会十多年来一直聘用众达律师事务所处理品牌和商标事务。这家律所是美国最大的律所之一,也从事打击工会的业务,但是:a) 维基媒体基金会似乎并未为此聘请他们;b) 这一合作关系早在当前这场风波之前很久就建立了。
https://news.ycombinator.com/item?id=49137144
Gmail’s search on mobile, where things come up in the quick results only to disappear when I complete the search, is a millstone about my neck.
mmargenot
Gmail在移动设备上的搜索功能,快速结果里明明出现了相关内容,但当我完成搜索时它们却消失了,这真是挂在我脖子上的沉重负担。
https://news.ycombinator.com/item?id=49131196
I’m not sure where this idea came from that farming was idle and only industry required constant work. Every hundred-year old book I’ve ever read that features farming includes themes of how the farmer’s work never ends and runs from dawn to dusk.
ip26
我不确定这种想法从何而来,认为农耕是清闲的,只有工业才需要持续劳作。我读过的每一本描写农耕的百年老书里,都有这样的主题:农民的活计永无止境,从黎明忙到黄昏。
https://news.ycombinator.com/item?id=49145034
Clicked like 5 pages and never found 1 code example.
Idk why languages don’t have their syntax in a sandbox front-and-center on the home page.
It’s like a video game site with zero screenshots or videos (also rampant).
New programming languages I want 2 things:
What does the syntax look like
Why would I use this language
Talk about the proof logic, show the syntax, thank you
cyanregiment
点了大概5页,一个代码示例都没找到。
不知道为啥语言不在首页正中放一个沙盒展示语法。
就像个视频游戏网站却没有任何截图或视频(这种还挺普遍)。
对于新编程语言,我只想要两件事:
语法长什么样
我为什么要用这门语言
讲证明逻辑、展示语法,谢谢。
https://news.ycombinator.com/item?id=49143529
Japan holds a huge amount of US treasuries, and I guess was considering a mass sell off to raise cash to defend the Yen.
US treasury bond yields are already dangerously high for the US and Japan selling treasuries would push yields up even higher, and could trigger more panic selling from others.
I guess this is Bessent’s scheme to try and kick that can down the road.
eigenspace
日本持有大量美国国债,我猜它原本在考虑大规模抛售以筹集现金来捍卫日元。
美国国债收益率对美方来说本已处于危险的高位,而日本抛售美债会进一步推高收益率,还可能引发更多人的恐慌性抛售。
我猜这就是贝森特的计策,试图把那个问题往后拖。
https://news.ycombinator.com/item?id=49129651
Step by step, Go is now learning the hard lessons every other language has learned over the last 20 years. The fact that despite their best efforts, their propositions look like everone else is a surprisingly refreshing affirmation of status quo.
DarkNova6
一步一步地,Go语言现在正在学习其他语言在过去20年里学到的艰难教训。尽管他们尽了最大努力,他们的提议看起来却与其他人无异,这出人意料地、令人耳目一新地肯定了现状。
https://news.ycombinator.com/item?id=49138385
I urge people to not read this. Once you do, you will see all documentation will as the flawed and confusing mess it is. Ignorance is bliss!
Hnrobert42
我恳请大家不要读这个。一旦你读了,你就会发现所有文档都是一团糟,漏洞百出又令人困惑。无知是福!
https://news.ycombinator.com/item?id=49131801
Why wouldn’t they be? No fiddling, super compact form factor, (usually) far superior energy efficiency, maybe phone apps, did I mention no fiddling?
Also, the form factor!
fuzzy2
为什么不会呢?没有折腾,超紧凑的形态,(通常)能效高得多,也许还有手机应用,我说过没有折腾吗?
还有,这个形态!
https://news.ycombinator.com/item?id=49130896
In “Four thousand weeks” by Oliver Burkeman, he breaks down that this obsession with action stems to the industrial revolution when it was decided that workers should sell their time for a living.
Before that we used to have task-oriented jobs, like milking the cows, which you cant do more than once in a while, or harvest the fields, which you can’t do until it’s ripe. In general humans are more evolved to this kind of work given our 200'000 years of task-based genetics versus 150 years of time controlling your action.
augment_me
在奥利弗·伯克曼的《四千周》一书中,他剖析了这种对行动的痴迷源于工业革命,当时人们决定工人应该靠出卖时间来谋生。
在那之前,我们做的是以任务为导向的工作,比如挤牛奶——你不能频繁地做这件事,或者收割庄稼——你必须等它成熟了才能收割。总的来说,人类更适合这种工作,因为我们有20万年的任务导向基因,相比之下,用时间控制你的行动只有150年的历史。