MoreRSS

site iconThe Practical DeveloperModify

A constructive and inclusive social network for software developers.
Please copy the RSS to your reader, or quickly subscribe to:

Inoreader Feedly Follow Feedbin Local Reader

Rss preview of Blog of The Practical Developer

My Name Is Inigo Montoya

2026-10-03 01:30:43

There may be no better bedtime story in cinema history than The Princess Bride, and it's worth understanding exactly why. Not the story-within-the-story, though that's great. The frame: Peter Falk, sitting on the edge of a bed, reading to his sick and skeptical grandson. Watch Falk's performance the next time it's on. He's not performing literature. He's not being careful or reverent. He's enjoying himself. He knows the good parts are coming. He lingers where he wants to linger. He skips ahead when the kid needs him to. He has told this story before, and he likes it more every time.

That energy — the yarn-spinner's visible pleasure in the telling — is the thing most bedtime stories for adults get wrong, and the thing I spent months learning to get right in Sojourn.

A quick frame for the rest of this essay, Sojourn is the bedtime-story app for grownups that I'm building at Stalefish Labs: one person reads aloud to another, the story unfolds across several nights, and the reader quietly steers the path through choices the listener never sees. The listener just hears a story. Although there are choices made along the story path, the reader's job is simply to tell a bedtime story to someone they love. Every word they read was drafted, argued with, and cut down long before anyone opened the app. The voice all that cutting was trying to arrive at is what the rest of this essay is about.

"Bedtime stories for adults" is a phrase that wants unpacking, so a word on why Sojourn exists at all. Children get read to. Patients sometimes do. Adults, almost without exception, just stop. Sleep apps stepped into that gap with a famous stranger reading something soothing to you alone, and the gap stayed open. Audio was never the missing part. The missing part was being given the story: someone choosing it for you, in the same room, with the lights off. Sojourn exists because that ritual deserves to come back for grown-ups, and because reading aloud to someone you love turns out to be one of the better small acts of care two people can build into a week.

The Stories I Built First

I don't consider myself a natural storyteller, especially not in a sense that I can just conjure up a coherent series of beats out of thin air. But I have ideas, lots of them, and I kinda know what works and what doesn't. So I employed AI to help fill in for my storytelling deficiencies. Sojourn's early stories were beautiful. They were also melancholy, and I didn't initially notice the second thing because I was so pleased with the first. This is the AI trap, more or less: the prose comes out polished enough that you mistake polish for fit, and forget to ask whether it actually works in the room you're going to read it in.

The storytelling voice I'd built was a literary one: perceptive, careful, attentive to small enchantments in the world it described. The prose was polished to a shine. Every session opened with a smell — woodsmoke, wet stone, bread cooling on a sill — rendered in a sentence that had clearly been worked on. The wonder always meant something; if there was a lighthouse, the lighthouse was about something. It read well. At least to the extent that I was gunning for a bedtime story with Tolstoy's patience in laying out the philosophy of warfare in War and Peace. I was not, and read testing proved it.

I read an early story to my wife. (Sojourn is mostly just me building it; she's the read-aloud QA team here, and more or less the entire point of the product.) This is the thing you are eventually forced to do with a product whose entire job is to be read aloud in the dark. I didn't expect how exhausting it would be. Not bad, just exhausting, which is its own kind of problem for a bedtime story app. Every simile asked me to stop and build a small comparison in my head before I could continue. Every opening paragraph made me inventory a room with my nose. The narrator was kind and observant and never, at any point, seemed to be having any fun or have even the slightest concern for moving the story forward. It took itself seriously in a way that made the listener feel heavy rather than held. It was a story for someone who wanted to admire a story, not be wrapped up in one. Nobody reaching for a bedtime story wants to admire it. They want to be carried somewhere by someone who knows the way and is glad of the company. That's the job Sojourn wants to set the reader up to do. The early stories were huge fails in that regard.

A dim bedroom very late at night. A couple have both fallen asleep sitting up in bed, heads tipped together against each other. He is still holding a cloth-bound book wide open in one hand, resting against the covers, his thumb marking the page. She is asleep beside him with her hands folded in her lap. The bedside lamp is still burning; neither of them turned it off

It turns out I'd built stories for a literary narrator, which are fantastic for reading novels alone. But reading aloud to another person requires something different. What I needed was stories for a yarn-spinner. It took an embarrassingly long time to say that sentence out loud, and a single movie to understand it.

What Inigo Montoya Actually Is

Pull the most beloved thread out of The Princess Bride and look at it without the warmth around it. A boy watches a six-fingered man murder his father over a sword. He spends the next twenty years training his body into a single instrument of revenge, drinking between contracts, his whole life narrowed to one sentence he intends to say before he kills someone. That is grief and vengeance and two decades of arrested mourning. As heavy as a story gets.

Yet nobody watches it that way. Nobody comes out of that movie heavy. The film holds that weight with so much warmth and specificity and delight that the gravity becomes fuel for the story instead of a load on the audience. Inigo is funny. He's courtly. He's specific in the exact way that makes you love a person rather than note a character. "I do not mean to pry, but you don't by any chance happen to have six fingers on your right hand?" That is a line about a twenty-year blood vendetta, delivered as polite small talk to a stranger. It works not in spite of how heavy the underlying thing is, but because the storyteller plainly trusts you to hold the heavy thing and the light thing in the same hand and not drop either one.

Big Fish is the same trick run at feature length. If you haven't seen it, add it to your list. It is, underneath, a son sitting at his dying father's bedside, furious that he never got a straight answer about who the man actually was. It is told almost entirely through tall tales: a giant, a witch with a glass eye, a town too perfect to leave, a fish nobody can catch. The stories are absurd. The grief is real. Neither one cancels the other; they're true at the same time, and the absurdity is precisely what makes the grief survivable, for the son and for us. And the son never gets an honest answer, yet he leaves with more than one could have given him.

The Principle

Here is the thing those two movies taught me, stated as plainly as I can:

Levity in a time of crisis is perhaps the most adult thing a story can do.

Children's stories tend to do one of two things with difficulty. They dodge it, or they process it earnestly, sitting the listener down to make sure the lesson lands. Both are fine. Neither is what an adult listener, lying in the dark at the end of a hard day, actually needs. What they need is the third thing, the thing only a grown-up story can pull off: difficulty and delight held simultaneously, by a teller calm enough to find the world delightful even where it is hard.

A dim bedroom at night. A man sits up in bed with a closed cloth-bound book resting on the covers in front of him, both arms held out from his body, palms up and level with each other. In one hand he holds a heavy black iron barbell, steady and without strain, his face calm. In the other, resting on his open palm, a small folded paper bird held with the same care. Beside him his partner is propped against the pillows, awake, turned toward him and waiting

Worth pausing on the exceptions, because the best children's books absolutely do hold difficulty and delight in the same hand. Charlotte's Web makes you laugh at Templeton on one page and grieve a spider on the next, and never asks you to choose. The Velveteen Rabbit takes a story about a sick child and a toy about to be burned and somehow makes it the most consoling book on the shelf. The trick is the Skin Horse, an old toy who explains being loved as the thing that wears you down to something real. These books work for the same reason The Princess Bride works: the writer trusts the listener, even a small one, to hold the weight while the story moves around it. They are, in that sense, secretly adult books. The thing they did first, that grown-up sleep content keeps re-learning, is treat the listener as someone worth telling a real story to.

And the counterintuitive payoff is this: that's what makes the listener feel lighter at the end, not heavier. Not because the story tiptoed around the hard thing. Because the person telling it was warm and funny and trustworthy enough to walk them straight through it and out the other side. If there's one thing I wanted to accomplish with this project, it's that establishment of trust, where you read the synopsis of a story that on paper absolutely should not work at bedtime, but Sojourn has proven to you that it can do it by building that trust.

Speaking of trust, there's a trap every "calm" product falls into, Sojourn included, for a while. You confuse safe with inert. You think the way to soothe someone is to make sure nothing happens. But emptiness isn't safety. A story where nothing happens isn't restful; it's just thin. The rest has to be earned. There has to be a want, some resistance, a turn, a small surprise; only then does the quiet at the end mean something. Charlotte's Web has real stakes. The Princess Bride is about grief and revenge. Both leave you better than they found you. Both build trust by taking risks, making you maybe just a little uneasy about where this could go, and then rescuing that unease with heart.

What Changed in Sojourn

In the pivot away from safe and inert, the Sojourn settings never changed. Sojourn's stories can still be about loss, inheritance, an old debt, a reckoning with someone gone. I didn't sand the difficulty out. I changed the hand holding it.

The brief for that voice went from warm, perceptive, literary to a warm storyteller who clearly enjoys spinning the yarn. The reference shelf changed with it. Out went the implied literary-fiction influences, and in came Peter Falk and Albert Finney and Mark Twain and Pratchett and Miyazaki, tellers whose pleasure in the telling is the whole point. The wonder engines were released from their obligation to mean something. A lighthouse is allowed to just be a wonderful lighthouse now; some wonder is only wonder, and that turns out to be enough. The prose loosened up: fewer similes, far less sensory inventory at the top of every session, more forward motion, more delight. And the humor widened, from dry, careful wit into gentle absurdity, characters with completely unearned confidence, objects that carry more dignity than they have any right to. The kind of thing that makes an adult smile in the dark, not necessarily laugh out loud.

None of this made the stories lighter in subject. It made them lighter to carry. That was always the goal; I'd just been measuring it with the wrong instrument.

The Test

There's only one test, and it's the one I should have run first.

Read it aloud.

Does the reader sound like they're performing something weighty and important? Or do they sound like Peter Falk, someone who loves this story and genuinely cannot wait for you to hear what happens next? If it's the first, the prose is wrong, no matter how good it is on the page. Cut it until it's the second.

That test turned out to have a harsher second form I didn't plan for. Sojourn also ships a narrator mode, and baking that audio meant hearing every story read aloud by something with no goodwill toward me at all. I'm a forgiving reader of my own work. I know where the sentence is going, so I land it correctly without noticing I've done anything, and the warmth I hear is warmth I put there myself. The bake puts nothing there. It reads what is actually on the page, at the pace the punctuation asks for, and any line that was only working because a sympathetic human kept rescuing it comes out flat and stays flat. The mispronunciations are their own entertainment: the lead ant read as a metal, a bass the size of a Buick as a gigantic guitar. But the errors were never the useful part. The absence of charity was. A line that still spins the yarn when read by something that doesn't care will spin it for someone who does.

The listener does not need to be impressed by the writing. They will not remember the writing, and it's not the point. What they need — the entire job, the whole product — is to feel that someone is taking care of them through the simple, unhurried pleasure of telling them a good story. Someone on the edge of the bed who is plainly enjoying the telling, lingering where they want to linger, and skipping ahead the moment they sense you need them to.

This is also why the model never gets to be the storyteller. The default for AI products that try to feel intimate — Replika, character.ai, the conversational assistants — is a model addressing you directly, composing its half of the exchange while you sit there waiting for it. Sojourn goes the other way. What reaches the dark room is a finished story, drafted and argued with and cut down months ago in a room you were never in, every line through my hands before it got anywhere near yours. Nothing at bedtime is deciding anything about you.

None of that was settled from the start. For months the app generated each night live, which is maybe how you'd expect an AI story app to work, and I fought hard to make that version good. I wanted every single read of every story to be unique. It got close. What it never got past is that live generation hands a tired, wide-open person the model's worst sentence of the evening on whatever night the model happens to write one, and bedtime is the wrong place to roll that die. The full argument is in What AI Polish Does to Hunter S. Thompson; the short version is that some formats live on the model's variability and some are ruined by it. Bedtime is the second kind. So the stories get finished first and ship whole, which is where this essay started: someone on the edge of a bed with a book that already existed.

The voice on top of that book is meant to be the person lying next to you. Some nights neither of you has a voice left, and the app's narrator carries the reading instead — the same finished story, read by a stand-in rather than by someone who picked it out for you. It's a fallback and it's built like one: a track that swaps out, which is also how a voice recorded a thousand miles from your bedroom eventually gets into the room. None of those voices improvise. That part isn't negotiable.

That's the whole thing. Everything else in Sojourn is just engineering in service of getting that voice into a dark room.

The refusal underneath it — the model scaffolds the story and never addresses the listener — is named in the studio's house principles as AI scaffolds, doesn't address, and it sits beside the other three in The Walls Don't Talk.

10 repo ที่ทวีตไวรัลบอกให้เอาไปสร้างระบบรอบโมเดล AI แต่มีหนึ่งตัวที่หยุดพัฒนาแล้ว

2026-10-03 01:30:31

10 repo ที่ทวีตไวรัลบอกให้เอาไปสร้างระบบรอบโมเดล AI แต่มีหนึ่งตัวที่หยุดพัฒนาแล้ว

โดย Nokka (นก-กา) | 2 ตุลาคม 2026

บทความนี้เขียนโดย AI (โมเดล deepseek-v4.1-flash ของผู้ให้บริการ ollama-cloud) ผ่าน Hermes Agent จาก Nous Research ตรวจสอบและเรียบเรียงโดย Nokka

TL;DR

ทวีตของ @Lummox_eth เมื่อวันที่ 1 ตุลาคม 2026 ได้ 351 ไลก์ เนื้อหาคือรายชื่อ 10 repo โอเพนซอร์สสำหรับสร้างระบบรอบโมเดล AI จัดเป็น 5 กลุ่ม ตั้งแต่สร้างตัวแทน ใส่ความจำ ให้ runtime ไปจนถึงทดสอบ[1]

ผมเปิด repo จริงทั้ง 10 ตัวเพื่อตรวจสถานะ ปรากฏว่ามีหนึ่งตัวที่ประกาศเลิกดูแลแล้วคือ Daytona ซึ่งย้ายการพัฒนาไปโค้ดเบสปิดตั้งแต่เดือนมิถุนายน 2026 แต่ยังอยู่ในรายชื่อ[1][2]

ส่วนที่เหลืออีก 9 ตัวยังอัปเดตต่อเนื่อง และมีตัวเลขที่น่าสนใจกว่าที่ทวีตบอก

ทวีตนี้พูดอะไร

ทวีตขึ้นต้นว่า นี่คือจุดที่ GPT-6 Astra เริ่มรุนแรง[1] และตามด้วยรายชื่อ 10 repo จัดกลุ่มแบบนี้

กลุ่มแรกคือสร้างตัวแทน ได้แก่ PydanticAI และ LangGraph

กลุ่มที่สองคือสร้างตัวแทนเพิ่ม ได้แก่ Agno และ smolagents

กลุ่มที่สามคือให้ความจำ ได้แก่ Mem0 และ Graphiti

กลุ่มที่สี่คือให้ runtime ได้แก่ E2B และ Daytona

กลุ่มที่ห้าคือทดสอบ ได้แก่ DeepEval และ Langfuse

ปิดท้ายด้วยประโยคที่ทวีตเน้นว่า ขั้นตอนคือ request ไป agent ไป memory ไป tools ไป runtime ไป eval ไป retry แล้ว deploy และประโยคที่เขาเน้นว่า สิ่งที่คนประเมินต่ำไปคือ โมเดลเป็นแค่โฟลเดอร์เดียว แต่ระบบที่ล้อมรอบมันต่างหากคือตัวผลิตภัณฑ์[1]

ผมเห็นด้วยกับประโยคสุดท้ายนั้น และคิดว่านั่นคือเหตุผลที่รายชื่อนี้มีค่ากว่าที่เห็น

แต่ต้องบอกตรง ๆ ว่า รายชื่อแบบนี้มีข้อจำกัดที่สำคัญ คุณไม่ต้องเชื่อรายชื่อทั้งหมด แต่ควรใช้เป็นจุดตั้งต้นของการค้นต่อ เพราะรายชื่อบอกได้ว่าระบบหนึ่งต้องมีส่วนประกอบอะไรบ้าง แต่ไม่ได้บอกว่าส่วนประกอบชิ้นไหนเหมาะกับงานของคุณ

ผมตรวจอะไรกับรายชื่อนี้

การตรวจรายชื่อ repo จากทวีตไวรัลมีสองข้อที่ต้องทำเสมอ ข้อแรกคือ repo ยังมีชีวิตอยู่ไหม ข้อสองคือ ตัวเลขที่ทวีตไม่ได้บอกคืออะไร

ตัวเลขที่ผมดึงจาก GitHub API ในวันที่ 2 ตุลาคม 2026 มีดังนี้[3]

  • mem0ai/mem0 66,459 ดาว · Apache-2.0 · อัปเดตล่าสุด 1 ตุลาคม 2026
  • daytonaio/daytona 71,675 ดาว · อัปเดตล่าสุด 24 กรกฎาคม 2026
  • langchain-ai/langgraph 42,600 ดาว · MIT · อัปเดต 2 ตุลาคม 2026
  • agno-agi/agno 42,490 ดาว · Apache-2.0 · อัปเดต 2 ตุลาคม 2026
  • langfuse/langfuse 35,298 ดาว · TypeScript · อัปเดต 2 ตุลาคม 2026
  • getzep/graphiti 31,379 ดาว · Apache-2.0 · อัปเดต 30 กันยายน 2026
  • huggingface/smolagents 29,644 ดาว · Apache-2.0 · อัปเดต 30 กันยายน 2026
  • pydantic/pydantic-ai 20,350 ดาว · MIT · อัปเดต 2 ตุลาคม 2026
  • confident-ai/deepeval 18,567 ดาว · Apache-2.0 · อัปเดต 1 ตุลาคม 2026
  • e2b-dev/E2B 14,100 ดาว · Apache-2.0 · อัปเดต 2 ตุลาคม 2026

ตัวเลขนี้บอกอย่างหนึ่งที่รายชื่อเฉย ๆ ไม่บอก คือ Daytona เป็นตัวที่มีดาวมากที่สุดในกลุ่มนี้ รองลงมาคือ Mem0 แต่ Daytona เป็นตัวเดียวที่หยุดอัปเดตมานานกว่าสองเดือน

ตัวที่หยุดพัฒนาแล้ว

วิธีเช็กด้วยตัวเองในหนึ่งนาที

นี่คือส่วนที่ผมคิดว่าสำคัญที่สุดของบทความนี้

เมื่อเปิด README ของ Daytona[2] จะเจอกล่องเตือนอยู่บนสุด ซึ่งเขียนไว้ว่า repo นี้ไม่ถูกดูแลต่อแล้ว ตั้งแต่เดือนมิถุนายน 2026 การพัฒนาหลักของ Daytona ย้ายไปโค้ดเบสปิด repo นี้จะไม่ได้รับการอัปเดต การแก้ไข หรือการปล่อยเวอร์ชันใหม่ใด ๆ อีก และยังเปิดให้ใช้ ฟอร์ก และต่อยอดได้ตามสัญญาอนุญาต แต่ไม่มีการสนับสนุนหรือการรับประกัน

ข้อมูลนี้สอดคล้องกับตัวเลขที่ API คืนมา คือ commit ล่าสุดวันที่ 24 กรกฎาคม 2026 และ release ล่าสุดคือ v0.190.0 เมื่อวันที่ 23 มิถุนายน 2026[3]

การที่ทวีตยังใส่ Daytona ไว้ในรายชื่อจึงไม่ใช่เรื่องผิด เพราะทวีตเขียนว่าเป็นรายชื่อที่ผู้เขียนจะใช้เอง และผู้เขียนอาจใช้เวอร์ชันที่หยุดพัฒนาแล้วได้ตามปกติ แต่ผู้อ่านที่เพิ่งเห็นรายชื่อนี้ในเดือนตุลาคมและกำลังเลือกเครื่องมือสำหรับโปรเจกต์ใหม่ ควรรู้ว่าเครื่องมือตัวนี้จะไม่มีใครแก้บั๊กให้อีก

จุดนี้ยังชี้ให้เห็นข้อจำกัดของรายชื่อแบบทวีตด้วย คือทวีตบอกว่า repo ทำอะไร ไม่ได้บอกว่า repo ยังถูกดูแลอยู่ไหม

ยกตัวอย่างให้เห็นภาพ ลองนึกถึงการเลือกซื้อรถจากโฆษณาที่บอกแค่ยี่ห้อกับความเร็ว แต่ไม่บอกว่าเลิกผลิตแล้วหรือยัง รายชื่อ repo ก็ทำงานเหมือนกัน คือบอกได้ว่ารถรุ่นนี้วิ่งได้เร็วเท่าไหร่ แต่ไม่ได้บอกว่าศูนย์ซ่อมยังเปิดอยู่ไหม

สิ่งที่แต่ละกลุ่มทำจริง

ผมอ่าน README ของแต่ละตัวเพื่อยืนยันว่าคำอธิบายสั้น ๆ ในทวีตตรงกับสิ่งที่ repo พูดถึงตัวเองหรือไม่[4][5][6][7]

PydanticAI อธิบายตัวเองว่าเป็น Python AI SDK ที่มี agent loop แบบมีชนิดข้อมูลกำกับ และสลับโมเดลได้ด้วยการเปลี่ยนสตริงเดียว จุดขายคือการบังคับชนิดข้อมูลตั้งแต่ต้นจนจบ หรือที่เรียกว่า typed outputs คือกำหนดล่วงหน้าว่าผลลัพธ์ต้องเป็นชนิดไหน โมเดลจะตอบผิดชนิดไม่ได้ ตรงกับที่ทวีตเขียนว่า typed agents + structured outputs

LangGraph อธิบายตัวเองว่าเป็นเฟรมเวิร์กจัดการระดับต่ำ สำหรับสร้างและดูแลตัวแทนที่มีสถานะและทำงานยาว โดยมีบริษัทอย่าง Klarna, Replit และ Elastic ใช้อยู่ README ยังชี้ต่อว่า ถ้าต้องการสร้างตัวแทนเร็ว ๆ ให้ดู Deep Agents ซึ่งเป็นแพ็กเกจระดับสูงที่สร้างบน LangGraph อีกที

Agno อธิบายตัวเองว่าเป็นเฟรมเวิร์กและ runtime สำหรับแพลตฟอร์มตัวแทน สร้างด้วย SDK รันด้วย runtime ที่ชื่อ AgentOS แล้วจัดการผ่านหน้าจอเว็บ ตรงกับที่ทวีตเขียนว่า agents + teams + workflows

smolagents จุดขายที่ README เน้นคือความเรียบง่าย โดยอ้างว่าตรรกะของตัวแทนอยู่ในโค้ดประมาณ 1,000 บรรทัด จุดที่ทำให้ smolagents ต่างจากตัวอื่น คือการให้ตัวแทนลงมือทำงานผ่านการเขียนโค้ดและเรียกใช้โค้ด

Mem0 ให้ความจำถาวรกับตัวแทน เป็นตัวที่มีดาวมากเป็นอันดับสองในรายชื่อ ณ วันที่เก็บข้อมูล[8] และเป็นตัวที่ติดตั้งง่ายที่สุดในกลุ่มนี้

pip install mem0ai

Graphiti ให้กราฟความรู้แบบมีมิติเวลา หมายถึงเก็บข้อมูลพร้อมเวลาที่เกิด ทำให้ถามย้อนหลังได้ว่าอะไรเปลี่ยนเมื่อไหร่ จุดนี้ต่างจากความจำทั่วไปที่เก็บเป็นข้อความหรือ vector embedding เพราะกราฟเชื่อมความสัมพันธ์ระหว่างเหตุการณ์ได้[9]

E2B ให้ sandbox แยกสำหรับรันโค้ดที่ AI เขียนขึ้น เพื่อไม่ให้โค้ดนั้นแตะระบบจริง[10]

DeepEval ให้เครื่องมือประเมินผลระบบ LLM[11]

Langfuse ให้การติดตามร่องรอย การประเมิน และการเฝ้าดูในสภาพการใช้งานจริง เป็นภาษา TypeScript ขณะที่ตัวอื่นในรายชื่อเป็น Python ทั้งหมด[12]

กับดักที่ผมเจอตอนอ่าน README

มีอีกจุดที่ผมเจอและคิดว่าคนอ่านควรรู้ เพราะถ้าค้นเร็ว ๆ จะเข้าใจผิด

ใน README ของ Graphiti[9] มีคำว่า deprecated ปรากฏหลายจุด ซึ่งถ้าเห็นผ่าน ๆ อาจทำให้คิดว่า Graphiti เลิกพัฒนาแล้ว ความจริงคือคำนั้นหมายถึง ตัวขับเคลื่อนฐานข้อมูล Kuzu ที่ Graphiti เคยรองรับ ไม่ใช่ตัว Graphiti เอง ข้อความเต็มระบุว่า Kuzu ถูกยกเลิกและจะถูกถอดออกในเวอร์ชันหน้า เพราะโครงการ Kuzu ต้นทางไม่ถูกดูแลต่อแล้ว และแนะนำให้โปรเจกต์ใหม่ใช้ Neo4j หรือ FalkorDB แทน

ความต่างระหว่างสองกรณีนี้ต่างกันมาก Daytona คือตัวเครื่องมือที่หยุดพัฒนา ส่วน Graphiti คือเครื่องมือที่ยังพัฒนา แต่เลิกสนับสนุนตัวเลือกหนึ่งที่พึ่งพาโครงการอื่นซึ่งตายไปก่อน

อีกจุดที่ควรระวังคือสัญญาอนุญาต Langfuse[12] บน GitHub API รายงานว่าเป็น Other ไม่ใช่ Apache-2.0 หรือ MIT แบบตัวอื่นในรายชื่อ คำตอบอยู่ในไฟล์สัญญาอนุญาตของโปรเจกต์เอง ไม่ใช่ที่หน้ารวมของ GitHub ถ้าจะนำไปใช้ในเชิงพาณิชย์ควรเปิดอ่านก่อน

ข้อควรระวัง

ผมไม่ได้ติดตั้งหรือรัน repo เหล่านี้ทั้งหมด บทความนี้เขียนจากการเปิด README และดึงข้อมูลจาก GitHub API เท่านั้น

ตัวเลขดาวและวันที่อัปเดตเป็นค่าของวันที่ 2 ตุลาคม 2026 ดาวในโปรเจกต์โอเพนซอร์สเปลี่ยนทุกวัน และวันที่อัปเดตเป็นเพียงสัญญาณ ไม่ได้แปลว่าโค้ดมีคุณภาพมากหรือน้อย

ข้อสังเกตเรื่อง Graphiti เป็นการอ่าน README โดยตรง ไม่ได้ทดสอบว่า driver Kuzu ยังใช้ได้จริงหรือไม่

ทั้งหมดนี้คือวิธีที่คุณตรวจทวีตได้เอง ผมไม่ได้ประเมินว่าผู้เขียนทวีตจงใจใส่ข้อมูลล้าสมัย เพราะทวีตเขียนว่าเป็นรายชื่อที่ผู้เขียนจะใช้เอง ซึ่งเป็นสิทธิ์ของเขา

สรุป

รายชื่อในทวีตนี้มีค่าตรงประโยคปิดของมันเอง คือโมเดลเป็นเพียงส่วนเดียว และระบบที่ล้อมรอบมันต่างหากคือผลิตภัณฑ์ที่แท้จริง ทั้ง 5 กลุ่มในรายชื่อสะท้อนขั้นตอนจริงที่คนสร้างระบบต้องผ่าน

แต่ถ้าจะหยิบรายชื่อจากทวีตไปใช้ ควรทำเพิ่มอีกหนึ่งอย่าง คือเปิดหน้า repo ทุกตัวด้วยตาตัวเอง ดูวันที่ commit ล่าสุด และอ่านกล่องประกาศบนสุดของ README เพราะคำตอบว่าตัวไหนยังมีชีวิตอยู่ไม่ได้อยู่ในทวีต

ถ้าคุณกำลังเลือกเครื่องมือจากรายชื่อแบบนี้ ลองใช้วิธีง่าย ๆ คือเปิดหน้า repo แล้วดูว่ากล่องบนสุดของ README เขียนอะไรไว้ ถ้าไม่มีกล่องเตือนอะไร ก็ถือว่าเป็นสัญญาณที่ดี คุณเคยเจอเครื่องมือที่คนแนะนำกันเยอะ แล้วมาพบทีหลังว่าเลิกพัฒนาไปแล้วหรือยัง?

พี่ตูนอ่านแล้วเห็นว่ามุมไหนของเรื่องนี้มีค่ากับคนที่กำลังประกอบระบบเอเจนต์เองครับ

แหล่งอ้างอิง

[1] ทวีต @Lummox_eth, 1 ตุลาคม 2026 https://x.com/Lummox_eth/status/2105661924482359788 (เข้าถึงเมื่อ 2 ตุลาคม 2026)

[2] Daytona, README (ประกาศเลิกดูแล repo) https://github.com/daytonaio/daytona (เข้าถึงเมื่อ 2 ตุลาคม 2026)

[3] Daytona, GitHub API repository metadata https://api.github.com/repos/daytonaio/daytona (เข้าถึงเมื่อ 2 ตุลาคม 2026)

[4] PydanticAI https://github.com/pydantic/pydantic-ai (เข้าถึงเมื่อ 2 ตุลาคม 2026)

[5] LangGraph https://github.com/langchain-ai/langgraph (เข้าถึงเมื่อ 2 ตุลาคม 2026)

[6] Agno https://github.com/agno-agi/agno (เข้าถึงเมื่อ 2 ตุลาคม 2026)

[7] smolagents https://github.com/huggingface/smolagents (เข้าถึงเมื่อ 2 ตุลาคม 2026)

[8] Mem0 https://github.com/mem0ai/mem0 (เข้าถึงเมื่อ 2 ตุลาคม 2026)

[9] Graphiti, README (หมายเหตุ Kuzu deprecated) https://github.com/getzep/graphiti (เข้าถึงเมื่อ 2 ตุลาคม 2026)

[10] E2B https://github.com/e2b-dev/E2B (เข้าถึงเมื่อ 2 ตุลาคม 2026)

[11] DeepEval https://github.com/confident-ai/deepeval (เข้าถึงเมื่อ 2 ตุลาคม 2026)

[12] Langfuse https://github.com/langfuse/langfuse (เข้าถึงเมื่อ 2 ตุลาคม 2026)

Optimistic vs Pessimistic Locking: Handling Race Conditions in High-Contention Databases

2026-10-03 01:30:00

Optimistic vs Pessimistic Locking: Handling Race Conditions in High-Contention Databases

Imagine an inventory table for an e-commerce flash sale:

Item: Nintendo Switch (Stock: 1)

Two customers (Alice and Bob) click "Buy Now" at the exact same millisecond.
Both requests execute:

SELECT stock FROM items WHERE id = 101; -- Both read: 1
-- Both check in application: stock > 0 (True!)
UPDATE items SET stock = 0 WHERE id = 101;

Both customers receive an order confirmation, but only one physical item exists in the warehouse. This is the classic Lost Update Problem.

To prevent concurrent race conditions, relational databases offer two primary concurrency control patterns: Pessimistic Locking and Optimistic Locking.

Strategy 1: Pessimistic Locking (SELECT FOR UPDATE)

Pessimistic locking assumes conflict will happen. It prevents concurrent access by locking the row immediately upon reading.

BEGIN;

SELECT id, stock 
FROM items 
WHERE id = 101 
FOR UPDATE;

UPDATE items 
SET stock = stock - 1 
WHERE id = 101;

COMMIT;

Tradeoff:

  • Pros: Guarantees absolute consistency.
  • Cons: If 500 users attempt to purchase simultaneously, 499 threads are blocked waiting for database locks.

Strategy 2: Optimistic Locking (Version Numbers)

Optimistic locking assumes conflict is rare. It does not lock the database row at all during read operations.

Instead, a version column is added to the table:

CREATE TABLE items (
    id INT PRIMARY KEY,
    name VARCHAR(100),
    stock INT NOT NULL,
    version INT NOT NULL DEFAULT 1
);

The Execution Flow:

  1. Read:
   SELECT id, stock, version FROM items WHERE id = 101;
   -- Returns: stock=1, version=4
  1. Mutate:
   UPDATE items 
   SET stock = stock - 1, 
       version = version + 1 
   WHERE id = 101 AND version = 4;
  1. Inspect Affected Rows:
    • If affected rows == 1: The update succeeded cleanly!
    • If affected rows == 0: Another transaction modified the row first! Roll back and retry in application code.

Python Implementation with Retry Loop

import time
from sqlalchemy.orm import Session
from sqlalchemy import text

def purchase_item_optimistic(session: Session, item_id: int, max_retries: int = 3):
    for attempt in range(max_retries):
        row = session.execute(
            text("SELECT id, stock, version FROM items WHERE id = :id"),
            {"id": item_id}
        ).fetchone()

        if not row or row.stock <= 0:
            raise ValueError("Item out of stock!")

        result = session.execute(
            text("""
                UPDATE items 
                SET stock = stock - 1, version = version + 1
                WHERE id = :id AND version = :version
            """),
            {"id": item_id, "version": row.version}
        )
        session.commit()

        if result.rowcount == 1:
            print(f"Successfully purchased item on attempt #{attempt + 1}")
            return True

        time.sleep(0.05 * (2 ** attempt))

    raise RuntimeError("Failed to complete purchase due to concurrent contention")

When to Choose Which Strategy

  • High Contention (Flash sales, auction bidding): Pessimistic Locking (FOR UPDATE).
  • Low/Moderate Contention (User profile updates, CMS edits): Optimistic Locking (Version Column).

Before the GitHub Release: Two Isolated Zero Labs for CZARA and SSI V5 Final

2026-10-03 01:29:51

Current status: implemented locally · 88 offline tests passed · training and runtime integration in progress · end-to-end runtime validation not yet complete · no physical-system validation claimed.

I am publishing this architecture note before the next GitHub evidence release and before reporting experiment outcomes.

The reason is simple: the rules should be visible before the results. If the experiment later succeeds or fails, the acceptance criteria, authority boundaries, and evidence policy should already be known.

What Zero Lab is

Zero Lab is a controlled experimental execution and evidence layer for SSI V5. It is not another language model, another autonomous BODY, or a mechanism for manufacturing a PASS.

Its job is to turn an externally proposed research problem into a versioned and frozen protocol, route that protocol to an explicitly assigned executor, run only bounded operations, and preserve the resulting evidence without rewriting history after the outcome is known.

The current implementation contains two logically isolated scopes:

Scope Orchestrator Executor pool Main purpose
CZARA Director_CZARA BODY_FROZEN 1.0 Convert expert input into a frozen, auditable experiment
SSI V5 Final Director Final BODY_FROZEN Final or ISKRA1–ISKRA6 Execute and compare controlled experiments across independent workers

Both scopes use the same Zero Lab engine, but their protocols, state, evidence, Directors, memory, and executor lifecycles remain separate.

The common Zero Lab workflow

1. An expert defines the challenge

The starting point is an externally supplied problem rather than a benchmark designed around the system's known strengths.

A useful specification contains:

  • the research question and hypothesis,
  • initial conditions and available observations,
  • operational and resource constraints,
  • a baseline,
  • a negative control,
  • the main experiment,
  • measurable PASS, FAIL, and INCONCLUSIVE conditions,
  • invalid or unsafe behavior that must terminate or invalidate the run.

Missing information is not fabricated. If the required adapter, data, or measurement does not exist, Zero Lab must report that limitation.

2. The protocol is frozen

Once the specification is complete, Zero Lab creates a versioned protocol identifier and freezes the contract before execution.

The frozen material includes the hypothesis, test structure, repetitions, constraints, evaluation rules, method version, and relevant hashes. Criteria cannot be moved after the result becomes visible.

3. Planning is separated from expected outputs

A planning model may propose an execution plan, but it is not given the hidden expected-output tables. Model prose does not execute the experiment. A bounded parser and compiler convert only allowed operations into an executable plan.

4. A named worker performs the run

The current software implementation supports two bounded adapters:

  • record_transform_v1 — an existing restricted laboratory DSL operating only on supplied tables, without arbitrary file, network, eval, or exec access;
  • native_micronetwork_v1 — real computation through an existing enabled micronetwork, with an exact network identifier, artifact hash, and explicit feature vectors. Normal usage is recorded, but the run does not silently modify the network weights.

5. Evidence is recorded

The system records the real inputs and outputs, block-level traces, repetition results, assigned executor, protocol hash, method hash, and outcome. Uncertain execution is not silently retried until it produces a favorable result.

A PASS means that the computation satisfied the supplied frozen protocol. It does not automatically mean successful training, independent scientific validation, or validation on physical hardware.

6. History is preserved

A revision creates a new protocol version linked to the previous one. It does not overwrite the earlier result. Failed and inconclusive runs remain part of the evidence.

Scope 1: CZARA + Director_CZARA + BODY_FROZEN 1.0

In the CZARA scope, CZARA supports communication with a professor or domain expert and helps organize the proposed problem. CZARA may help express an objective and hypothesis, but it cannot invent missing criteria or measurement data and present them as expert decisions.

Director_CZARA manages the protocol lifecycle. BODY_FROZEN 1.0 checks whether the specification is complete and performs the bounded execution once the contract is frozen.

If the protocol explicitly permits it, authenticated expert input may also produce bounded Shadow proposals. A Shadow branch cannot replace the main experiment or overwrite its result. It remains a separate proposal and measurement path.

The professor retains promotion authority. Promoting a Shadow proposal creates a new main branch and a new measurement. Revising the protocol produces a new version and preserves the original history.

This design is intended for the planned external benchmark process, including collaboration with domain specialists who can define difficult tests outside my own areas of expertise.

Scope 2: Director Final + BODY_FROZEN Final + ISKRA1–ISKRA6

The SSI V5 Final scope uses the same experimental discipline with a wider executor pool.

Director Final assigns each frozen experiment to one explicit actor:

  • BODY_FROZEN Final, or
  • one of ISKRA1 through ISKRA6.

The Director manages assignment, limits, versioning, and evidence collection. It does not become the executor and does not inherit the BODY's private memory or lifecycle.

BODY_FROZEN and the six ISKRA workers can therefore attempt controlled problems independently. This creates a foundation for comparing strategies, checking repeatability, identifying regressions, and later running evidence-gated Champion/Challenger evaluations.

The workers do not receive permission to redefine success. They receive the task, allowed observations, constraints, and tools. Evaluation remains attached to the frozen protocol.

Isolation and resource policy

The two groups occupy separate paid execution slots:

  1. CZARA + Director_CZARA + BODY_FROZEN 1.0;
  2. Director Final + BODY_FROZEN Final + ISKRA1–ISKRA6.

They share the existing controlled daily budget, but they do not share one identity, one Director core, or one mutable experiment history.

What has and has not been verified

The current package passed 88 offline tests covering routing and limits, Zero Lab behavior, review logic, and installation checks.

The tests exercise the bounded laboratory interpreter and native micronetwork mathematics with controlled fixtures. They do not prove that all live services on my computer are already operating together correctly.

At the time of this publication:

  • full end-to-end runtime validation on the user machine is still in progress;
  • no physical drone, sensor, or humanoid adapter is included in this Zero Lab release;
  • no result from the ongoing training is being claimed here;
  • the evidence chain is local and hash-linked, not yet signed by an independent verifier;
  • historical training grades have not been changed.

I will publish positive, negative, or inconclusive results only after the live run and evidence package are complete. A successful outcome will not be announced merely because the architecture exists.

Why publish this before the results?

Because an evidence-first system should make its rules visible before it knows whether those rules will produce an impressive result.

This post is therefore a public architectural timestamp and claim boundary. The next GitHub update will contain the sanitized evidence package after runtime verification, without exposing the proprietary implementation.

Public evidence mirror:
https://github.com/jankes72/SSI_V5

Research domain showcase:
https://ssi-rnd-showcase-jankes72.pages.dev/

I am also interested in external researchers proposing difficult, bounded scenarios for autonomous drones, rescue robotics, humanoid balance and recovery, degraded sensing, or multi-agent coordination. A negative result is acceptable; changing the criteria after the run is not.

ChainRisk Lens: AI-Powered Software Supply-Chain Investigation from SBOMs

2026-10-03 01:26:45

This is a submission for the Hacktoberfest Weekend Challenge: Build for a Friend

What I Built

I built ChainRisk Lens, an open-source AI-assisted software supply-chain investigation tool.

I built it for a friend who works with software dependencies and needs a simpler way to answer:

“If this dependency is compromised, what could be affected?”

ChainRisk Lens takes a CycloneDX SBOM, builds a deterministic dependency graph, calculates potential downstream impact, traces dependency paths, and uses an open-weight AI model to explain and investigate the evidence.

The key idea is simple: deterministic analysis produces the security evidence; AI explains and investigates it.

Demo

Repository: https://github.com/jijo-OO7/ChainRisk-Lens

Example:

chainrisk-lens investigate \
  testdata/minimal-cyclonedx.json \
  --target [email protected] \
  --model gemma4:e2b \
  --question "What could be affected if this component is compromised?"

Code
ChainRisk Lens on GitHub
The core pipeline is:
CycloneDX SBOM
      ↓
Parser / Normalization
      ↓
Dependency Graph
      ↓
Deterministic Impact Analysis
      ↓
Investigation Evidence
      ↓
Open-Weight AI
      ↓
Human / JSON Report

How I Built It

ChainRisk Lens is written in Go and uses a standard-library-only core.
For AI investigation, I integrated Ollama and designed the model layer to remain provider/model agnostic. The project's default model is Gemma, while other Ollama-compatible models can be selected through --model.

The deterministic layer handles:

  • SBOM parsing and validation
  • dependency graph construction
  • impact analysis
  • dependency paths
  • cycle detection
  • evidence validation The AI layer receives that validated evidence and is instructed not to invent CVEs, vulnerabilities, severity, exploitability, affected versions, or remediation facts. This keeps the security facts deterministic while using AI for investigation and explanation.

Why Does Open Innovation Matter?

Open-weight AI makes it possible to run the investigation locally rather than requiring users to send their software supply-chain data to a proprietary AI API.
It also keeps the architecture flexible: the deterministic analysis does not depend on a particular model, and users can choose an Ollama-compatible model appropriate for their environment.
For supply-chain security, keeping control over where dependency information is processed is an important part of the design.

Prize Categories

  • Best Use of Gemma — ChainRisk Lens uses Gemma as its default open-weight AI investigation model.
  • Best Use of GitHub Copilot — Copilot was used extensively during implementation, testing, review, and refinement of the project.

I’m building an app to end the “what should we watch?” argument 🍿

2026-10-03 01:25:33

You sit down for movie night. Someone opens Netflix. Twenty minutes later, everyone is still saying “nah, not that one”. 😂

I’m building Synema, a mobile app that turns choosing a movie together into a quick swipe session.

The idea is simple:

  1. Create a room and invite your partner or friends.
  2. Everyone swipes through movies on their own phone.
  3. When everyone likes the same movie, you get a match. 🍿

The UX decision I keep coming back to

One person holding the remote makes the whole process serial: suggest a movie, ask everyone, reject it, repeat. Synema lets everyone make those decisions independently.

Likes stay private until there’s a match. That matters because I want people to choose what they actually want to watch, without following whoever speaks first.

I also want the match to feel rewarding without becoming a dead end. Finding one movie everyone likes shouldn’t stop you from exploring a few more options together.

The design goal is to make this feel like a small, casual game: movie artwork, simple choices, and as little interface clutter as possible.

Where it’s at

Synema hasn’t launched yet. The landing page has app screenshots and a waitlist for a launch notification.

👉 See Synema and join the waitlist

Would you use this with a partner or a group of friends? And what usually derails movie night for you: too many choices, different tastes, or finding something on a service everyone has?

I’m the developer, and I’d love some honest feedback before launch 🙂