NewsJack
这是 AlexAImaginator 公开的项目资料。本站按原始来源展示项目信息,用中文说明适用场景和阅读边界;项目名、源帖与代码保持原样,便于逐项核对。
这条案例记录了什么
Agent 与模型路由、Agent 评估
根据任务状态选择一条已定义的处理路径。
第三方转述
原始来源:Secondary X breakdown and GitHub repository。作者自述,本站未独立复现。
AlexAImaginator
记录日期:2026-09-18。日期与身份应以原始资料为准。
怎样核对这个项目
- 先打开原始来源,确认作者、日期与 Jev 在项目中的具体用途。
- 如果提供仓库,再检查代码、运行要求和许可证;仓库存在不代表本站已经运行成功。
- 对速度、成本、准确率和规模数字,查看原文的任务、环境和计算口径。
- 路由候选是否完整。
- 模糊请求的回退分支。
- 完整任务成本与结果。
- 是否有可观察的正确答案。
原始文字与技术细节
以下内容保留原语言,供核对事实。中文页的场景说明是阅读提示,不是逐句翻译或实测结论。
展开英文项目摘要与原帖
项目摘要
The cited NewsJack workflow uses Jev to make bounded triage decisions over incoming news and sends the selected opportunities to brand-specific agents instead of asking a frontier model to process everything.
来源原文
Explanation of what Jev is and how the whole thing works — for myself and for anyone who still doesn't get it 🙂 Jev (made by TypeSafe) is not an LLM and not ChatGPT. It doesn't return text. You give it context + a question with defined options (yes/no, pick from a list, score 0–4) and it returns a calibrated decision — a number with a probability, not a sentence. You define the answer space, it just distributes probability over it. That's why it's fast (~0.3 s) and ~400x cheaper than an LLM, and why it fits the "boring" jobs — triage, classification, routing, guardrails — where you want a decision, not an essay. The expensive model (or a human) only runs on the small bit Jev lets through the gate. Elvis Sun showed it on newsjacking: 384 headlines × 15 brands in 25 s for $0.19. # What Elvis Sun Did With Jev — A Practical Breakdown Source post: https://x.com/elvissun/status/2100951347080421409 (by Elvis Sun, @elvissun) Repo: https://github.com/elvisun/newsjack (demo in demos/news-desk-dealer/) Jev is made by TypeSafe (https://typesafe.ai) — the AI lab that built it. Written for a non-PR audience — explains the pattern in plain terms. ## The claim in one line "Jev is INSANE. In 24.9 seconds it read 384 news from this morning and told 15 brands which stories to hop onto today, for $0.19. Claude Opus 5, running on the same feed at the same time, got through 4/384 and cost $0.77. Per headline that is ~390x cheaper, and the answer comes back before you finish reading the headline yourself." He then open-sourced the whole thing as http://newsjack.sh — a set of skills that turn a coding agent (Claude Code, Codex, etc.) into a PR team, with the Jev demo sitting in the demos/ folder. ## The context: what "newsjacking" is Newsjacking is a PR tactic: when a big story breaks in your space, you attach your brand to it while the story is still hot. The window is often hours. A human PR team can't read 400 headlines a morning, score each one, and match it to 15 client brands in time. That's the problem Jev is solving here. The task, precisely: 1. Take 384 fresh headlines from Google News (fetched via RSS, committed to the repo). 2. For 15 illustrative brands (one per industry — tech, finance, health, consumer, etc.), figure out which stories each brand should "hop onto" today. 3. Do it fast enough to act on before the wave breaks. ## How Jev was actually used — two question layers Elvis did not ask Jev to "write a PR pitch." That would be an LLM job. Instead he asked Jev a set of structured, closed-space questions in two layers. This is the whole point of the model: it returns typed decisions with probabilities, not prose. Layer A — "Judge the story" (one call per headline, 9 questions in a single API call): 1. is_news (yes/no probability) — Is this a real news item, or a product page / listicle / SEO filler? 2. desk (pick one) — Which newsroom desk owns it: tech / business / finance / health / science / policy / consumer / culture / world 3. story_type (pick one) — breaking_event / announcement / data_report / regulation_policy / funding_deal / personnel_move / opinion_analysis / trend_feature / incident_crisis 4. magnitude (score 0–4) — 0 = niche, few outlets ... 4 = historic, defines the week 5. velocity (score 0–4) — 0 = static, one outlet ... 4 = everywhere at once, viral 6. novelty (score 0–4) — 0 = same story every cycle ... 4 = a first, no precedent 7. window (score 0–4) — 0 = a month of runway ... 4 = live breaking, 30-minute window 8. heat (score 0–4) — 0 = nobody arguing ... 4 = coverage turned into backlash 9. risk (score 0–4) — 0 = clean ... 4 = rides directly on tragedy, kill-switch territory Each score criterion carries a short description (e.g. window: "30min: live breaking, instant reaction only"), which is what Jev uses to calibrate the number. This is the part people miss: you define the scale, Jev fills in a probability-weighted number on it. After the call, plain code (no model) combines the scores into a newsworthiness score out of 10 using fixed weights: newsworthiness = magnitude×0.25 + velocity×0.25 + novelty×0.15 + window×0.15 (+ standing×0.2 if Layer B ran) The weights are deterministic code, not a model decision. The model only ever produced the 0–4 axis numbers. Layer B — "Match the brand" (one call per headline × brand pair, 6 questions; 384 × 15 = up to 5,760 brand evaluations): 1. standing (score 0–4) — Does the company have a real reason to be quoted: none / observer / adjacent / category / direct 2. journalist_shape (score 0–4) — Would a named reporter on a real beat want this company's take: no reporter / stretch / plausible / likely / obvious 3. bridge_type (pick one) — Which angle gives the company an honest way in: expert_reaction / data_contrast / contrarian / explainer / customer_proof / none 4. tier (pick one) — pitch_ready / big_story / watch 5. action (pick one) — ride (act now) / wait / skip / avoid 6. off_policy (yes/no probability) — Would pitching this violate the company's stated exclusions? So the final output per (headline, brand) pair is: a newsworthiness score, a standing score, a recommended angle, a tier, and an action — all typed, all bounded, all with probabilities underneath. ## The actual numbers (from the post + the demo README) - 384 headlines, 15 brands, morning of Sept 18, 2026. - Jev: 24.9 s total, $0.19. The demo README reports live measurements: Layer A ~190 ms per call, Layer B (90 questions, ~10k input tokens) ~320 ms per call, ~1,200 judgments/sec with 8 parallel workers. - Claude Opus 5 on the same feed: 4–6 s per headline for Layer A alone. The run was stopped the moment Jev finished so no extra tokens were burned. By that point it had completed 4/384 headlines and already cost $0.77. A full Opus pass over all 398 headlines is estimated at ~$25. - Per-headline cost: ~390x cheaper with Jev. - Elvis's follow-up ("Jev Mode") shipped this as a real media-monitoring feature: ~200 signals in 8 seconds for about a cent, "zero dropped stories" in their eval. ## Why this is the right tool for this job (the pattern) This is not "Jev is a cheap GPT." It is a specific architectural pattern: 1. Closed answer spaces. Every question has a finite, pre-defined set of answers. Jev cannot hallucinate a "desk" that isn't in the list; it can only distribute probability over the options you gave it. 2. Calibrated probabilities. The numbers are trained to match outcomes. A window score of 3.7 means "high probability this is a ~4-hour window," not "the model feels it's urgent." 3. Batch-friendly. Many questions per call, many calls in parallel. The cost is per input token, so you can score thousands of items for cents. 4. The expensive step only runs on survivors. This is the "gate" pattern: Jev screens all 384 headlines first; only the stories that clear the bar get any further (human review, or a real LLM drafting the actual pitch). In Elvis's real pipeline that is literally a skill called /relevance-coarse-filter — "cheap high-recall first pass that throws out obvious junk before anything expensive runs." The division of labor: - Jev = the fast, cheap judge that sorts and scores everything. - Deterministic code = the weights, thresholds, and final newsworthiness math. - LLM / human = only what survives the gate (drafting pitches, creative angles). ## The honest caveats - The headline demo in the repo ships in mock mode by default (keyword heuristics + a seeded random number generator, no API keys). The live Jev/Opus engines are wired but off until you add keys and set ?engine=live. The numbers in the post are from his live run, which the README documents as measured on 2026-09-18. - Jev's scores are probability-weighted floats rounded to levels — a "4" isn't "certainly historic," it's "mostly probability mass on the historic bucket." - "Zero dropped stories" is an eval claim by the author, not an independent benchmark. - Jev does not write the pitch. It decides which story which brand should act on, and with what angle. The actual copy is a separate, later step. ## The takeaway Jev is a decision model, not a chat model. Its sweet spot is exactly what Elvis built: high-volume, time-sensitive, closed-space judgment — classify 400 headlines, score them on 10 bounded axes, match them to 15 brands, do it in 25 seconds for $0.19 — where a frontier LLM would be 100x slower, 400x more expensive, and would still give you prose you have to parse. The pattern transfers anywhere you have many items, a fixed set of categories, a time pressure, and a budget: triage, routing, filtering, monitoring, guardrails — the "boring 80%" of a pipeline where an LLM is overkill. #Jev #TypeSafe #SystemOne #AIAgents #Newsjacking #LLM #AI
原记录的限制
- The reviewed source is secondary and does not establish the author's ground-truth quality.
- News relevance is time-sensitive; a throughput result is not an editorial-quality benchmark.
- The expensive synthesis or human review step remains necessary for publishable output.