AIHOT Open-Sourced: Build Your Own Industry Hot-Topic Site by Swapping Sources
The article introduces AIHOT, an open-source AI-powered news aggregation system that automatically collects, filters, scores, clusters, and publishes daily briefings; its methodology (prompts, scoring, thresholds) is fully open, allowing anyone to adapt it to their industry by changing sources and calibration data.
What Is AIHOT?
AIHOT is a complete, production-grade website that turns a list of information sources into a curated hot-topic site. You provide sources (RSS, web lists, JSON APIs, X accounts, WeChat public accounts, or custom scripts), and the system automatically fetches content daily, pre-filters with an LLM, runs two independent scoring passes against the same criteria, clusters related reports into single events, ranks events by a heat algorithm that resists manipulation, and publishes a daily briefing at 8 AM plus weekly and monthly digests. The repository contains the full stack: frontend, backend, selection pipeline, clustering and heat algorithms, and — most importantly — every prompt and acceptance threshold in plain text.
Performance is notable: median page response is 10 ms, with 95% of requests under 50 ms, running on Node.js 24 and PostgreSQL without any external BaaS.
The Six-Step Selection Pipeline
Each piece of content passes through six stages:
Collection — pull raw items from configured sources.
Pre-filter — an LLM does a quick relevance pass.
First scoring — the same item is scored against the selection criteria.
Second scoring — an independent second pass with the identical prompt; both must clear the threshold.
Writing — accepted items get a Chinese title and summary.
Clustering & publication — vector search over the last two weeks finds candidate duplicates; an LLM judges whether they are the same event, a follow-up, or distinct. Follow-ups attach to the parent event, which then receives a synthesized overview. Heat is computed per event (not per article): each unique source counts once per 48 hours, with a 24-hour half-life, preventing volume spam from inflating rankings.
Three Key Design Decisions
1. Independent double scoring with the same rubric
LLM scoring is non-deterministic. AIHOT runs the identical prompt twice; only items that pass both passes enter the pool. Source-tier thresholds (official vs. self-media) are layered on top. Every prompt and threshold is exposed in the repo for line-by-line editing.
2. Event-level clustering
Multiple outlets covering the same story are merged. The system first uses title/summary embeddings to retrieve candidates from the prior two weeks, then asks an LLM to classify the relationship (same event, follow-up, or different). Uncertain cases are re-checked with a second model. The event page shows a consolidated narrative.
3. Heat by event, not by article
Each distinct source contributes at most once per 48-hour window; heat decays by 50% every 24 hours. A single publisher pushing ten articles or a bot farm reposting en masse cannot game the leaderboard. The author notes this anti-gaming rule is the "lifeline" of a site that decides daily what deserves attention.
Why Open-Source?
Professionals from law, HR, finance, precious metals, etc., asked the author (Kazik) to build industry-specific versions. His README answer: "I can't. I don't know your industry, which sources matter, or what counts as a hot topic. But you do. Since I can't serve everyone, I'm handing you the spark."
The core open-sourced asset is not the few thousand lines of TypeScript but the complete methodology: prompt texts, scoring rubrics, acceptance thresholds, and calibration procedures. The code is intentionally generic; all industry-specific logic lives under industry/, especially industry/prompts/selection-score.md.
Customization and Calibration
Adapting the site is designed to be trivial. The recommended workflow:
Feed the repo to an AI coding assistant (Claude Code, Codex) with a prompt like:
请读 AGENTS.md 和 docs/customize.md,把这个站改成「法律」行业的热点站。 我关心的是:……(你想盯哪些信源,什么消息重要、什么不重要,越具体越好)Run the calibration loop: annotate a few hundred items, then execute scripts/eval-selection.ts to run SelectBench against your samples. Iterate the rubric until precision/recall meet your bar. Selection quality becomes measurable, not mystical.
Boundaries are clear: the repo ships with 18 overseas AI sources as demos, not the real AIHOT source list or operational data. The AIHOT name and logo are excluded from the MIT license — "swap in your own brand, and it becomes your site."
Deployment
One Docker-enabled machine plus an OpenAI-compatible API key (DeepSeek, Qwen, Zhipu, etc.) is enough:
git clone https://github.com/KKKKhazix/AIHOT.git myhot
cd myhot
node scripts/init-env.ts --llm-key <你的模型APIKey>
docker compose up -d --buildVisit http://localhost:3000; admin panel at /admin. Content starts appearing within minutes; the first full batch finishes in ~30 minutes.
Related Projects: Four Delivery Forms
The article positions AIHOT alongside three other open-source tools, forming a spectrum of delivery shapes for AI-processed information:
AIHOT (this project) — build a site for an audience .
Horizon (9,500+ stars) — push a personal bilingual briefing to WeChat, Lark, or email. Includes a built-in "AI creator" profile that breaks each story into event, timeliness hook, and angle suggestions.
last30days-skill-cn (1,800+ stars) — ask an agent "what has the AI circle discussed in the last 30 days?" It searches Weibo, Xiaohongshu, Bilibili, Zhihu, Douyin, etc. (8 platforms), producing a cited research report. Zero API keys except WeChat.
MediaCrawler (66,000 stars) — raw data layer for seven Chinese platforms using Playwright to handle login state and signatures; no JS reverse-engineering required. License permits learning/research only, no commercial use.
newspaper (15,000 stars) — the veteran extraction library (since 2013) that turns a URL into clean text, author, date, images, keywords. AIHOT, Horizon, and similar pipelines all rely on this step.
The choice depends on who consumes the output: a public audience, yourself, an agent, or your own custom pipeline.
Conclusion
Traditional open-source releases a library, framework, or component. Kazik has open-sourced an entire "how to run a media operation" playbook: source tiering, double-blind scoring, manipulation-resistant heat, and sample-driven calibration. The technical half is shipped; the domain-knowledge half is left for each industry's experts to fill in.
Signed-in readers can open the original source through BestHub's protected redirect.
This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactand we will review it promptly.
Geek Labs
Daily shares of interesting GitHub open-source projects. AI tools, automation gems, technical tutorials, open-source inspiration.
How this landed with the community
Was this worth your time?
0 Comments
Thoughtful readers leave field notes, pushback, and hard-won operational detail here.
