AIHOT Open-Sourced: Build Your Own Industry Hot-Topic Site by Swapping Sources

The article introduces AIHOT, an open-source AI-powered news aggregation system that automatically collects, filters, scores, clusters, and publishes daily briefings; its methodology (prompts, scoring, thresholds) is fully open, allowing anyone to adapt it to their industry by changing sources and calibration data.

Geek Labs
Geek Labs
Geek Labs
AIHOT Open-Sourced: Build Your Own Industry Hot-Topic Site by Swapping Sources

What Is AIHOT?

AIHOT is a complete, production-grade website that turns a list of information sources into a curated hot-topic site. You provide sources (RSS, web lists, JSON APIs, X accounts, WeChat public accounts, or custom scripts), and the system automatically fetches content daily, pre-filters with an LLM, runs two independent scoring passes against the same criteria, clusters related reports into single events, ranks events by a heat algorithm that resists manipulation, and publishes a daily briefing at 8 AM plus weekly and monthly digests. The repository contains the full stack: frontend, backend, selection pipeline, clustering and heat algorithms, and — most importantly — every prompt and acceptance threshold in plain text.

Performance is notable: median page response is 10 ms, with 95% of requests under 50 ms, running on Node.js 24 and PostgreSQL without any external BaaS.

The Six-Step Selection Pipeline

Each piece of content passes through six stages:

Collection — pull raw items from configured sources.

Pre-filter — an LLM does a quick relevance pass.

First scoring — the same item is scored against the selection criteria.

Second scoring — an independent second pass with the identical prompt; both must clear the threshold.

Writing — accepted items get a Chinese title and summary.

Clustering & publication — vector search over the last two weeks finds candidate duplicates; an LLM judges whether they are the same event, a follow-up, or distinct. Follow-ups attach to the parent event, which then receives a synthesized overview. Heat is computed per event (not per article): each unique source counts once per 48 hours, with a 24-hour half-life, preventing volume spam from inflating rankings.

AIHOT 六步流水线
AIHOT 六步流水线

Three Key Design Decisions

1. Independent double scoring with the same rubric

LLM scoring is non-deterministic. AIHOT runs the identical prompt twice; only items that pass both passes enter the pool. Source-tier thresholds (official vs. self-media) are layered on top. Every prompt and threshold is exposed in the repo for line-by-line editing.

2. Event-level clustering

Multiple outlets covering the same story are merged. The system first uses title/summary embeddings to retrieve candidates from the prior two weeks, then asks an LLM to classify the relationship (same event, follow-up, or different). Uncertain cases are re-checked with a second model. The event page shows a consolidated narrative.

五个来源的报道聚成一个事件,进入热点榜
五个来源的报道聚成一个事件,进入热点榜

3. Heat by event, not by article

Each distinct source contributes at most once per 48-hour window; heat decays by 50% every 24 hours. A single publisher pushing ten articles or a bot farm reposting en masse cannot game the leaderboard. The author notes this anti-gaming rule is the "lifeline" of a site that decides daily what deserves attention.

Why Open-Source?

Professionals from law, HR, finance, precious metals, etc., asked the author (Kazik) to build industry-specific versions. His README answer: "I can't. I don't know your industry, which sources matter, or what counts as a hot topic. But you do. Since I can't serve everyone, I'm handing you the spark."

The core open-sourced asset is not the few thousand lines of TypeScript but the complete methodology: prompt texts, scoring rubrics, acceptance thresholds, and calibration procedures. The code is intentionally generic; all industry-specific logic lives under industry/, especially industry/prompts/selection-score.md.

Customization and Calibration

Adapting the site is designed to be trivial. The recommended workflow:

Feed the repo to an AI coding assistant (Claude Code, Codex) with a prompt like:

请读 AGENTS.md 和 docs/customize.md,把这个站改成「法律」行业的热点站。 我关心的是:……(你想盯哪些信源,什么消息重要、什么不重要,越具体越好)

Run the calibration loop: annotate a few hundred items, then execute scripts/eval-selection.ts to run SelectBench against your samples. Iterate the rubric until precision/recall meet your bar. Selection quality becomes measurable, not mystical.

Boundaries are clear: the repo ships with 18 overseas AI sources as demos, not the real AIHOT source list or operational data. The AIHOT name and logo are excluded from the MIT license — "swap in your own brand, and it becomes your site."

Deployment

One Docker-enabled machine plus an OpenAI-compatible API key (DeepSeek, Qwen, Zhipu, etc.) is enough:

git clone https://github.com/KKKKhazix/AIHOT.git myhot
cd myhot
node scripts/init-env.ts --llm-key <你的模型APIKey>
docker compose up -d --build

Visit http://localhost:3000; admin panel at /admin. Content starts appearing within minutes; the first full batch finishes in ~30 minutes.

Related Projects: Four Delivery Forms

The article positions AIHOT alongside three other open-source tools, forming a spectrum of delivery shapes for AI-processed information:

AIHOT (this project) — build a site for an audience .

Horizon (9,500+ stars) — push a personal bilingual briefing to WeChat, Lark, or email. Includes a built-in "AI creator" profile that breaks each story into event, timeliness hook, and angle suggestions.

last30days-skill-cn (1,800+ stars) — ask an agent "what has the AI circle discussed in the last 30 days?" It searches Weibo, Xiaohongshu, Bilibili, Zhihu, Douyin, etc. (8 platforms), producing a cited research report. Zero API keys except WeChat.

MediaCrawler (66,000 stars) — raw data layer for seven Chinese platforms using Playwright to handle login state and signatures; no JS reverse-engineering required. License permits learning/research only, no commercial use.

newspaper (15,000 stars) — the veteran extraction library (since 2013) that turns a URL into clean text, author, date, images, keywords. AIHOT, Horizon, and similar pipelines all rely on this step.

AI 信息处理的四种交付形态
AI 信息处理的四种交付形态

The choice depends on who consumes the output: a public audience, yourself, an agent, or your own custom pipeline.

Conclusion

Traditional open-source releases a library, framework, or component. Kazik has open-sourced an entire "how to run a media operation" playbook: source tiering, double-blind scoring, manipulation-resistant heat, and sample-driven calibration. The technical half is shipped; the domain-knowledge half is left for each industry's experts to fill in.

Original Source

Signed-in readers can open the original source through BestHub's protected redirect.

Sign in to view source
Republication Notice

This article has been distilled and summarized from source material, then republished for learning and reference. If you believe it infringes your rights, please contactadmin@besthub.devand we will review it promptly.

Node.jsOpen sourcePostgreSQLLLM promptingAI news aggregationindustry-specific newsnews clusteringSelectBench
Geek Labs
Written by

Geek Labs

Daily shares of interesting GitHub open-source projects. AI tools, automation gems, technical tutorials, open-source inspiration.

0 followers
Reader feedback

How this landed with the community

Sign in to like

Rate this article

Was this worth your time?

Sign in to rate
Discussion

0 Comments

Thoughtful readers leave field notes, pushback, and hard-won operational detail here.