SKILL.md
readonly只读
name
data-scraper-agent
description
构建一个完全自动化的AI驱动数据收集代理,适用于任何公共来源——招聘网站、价格、新闻、GitHub、体育等。按计划运行,使用免费LLM(Gemini Flash)丰富数据,将结果存储在Notion/Sheets/Supabase中,并从用户反馈中学习。在GitHub Actions上100%免费运行。当用户希望自动监控、收集或跟踪任何公共数据时使用。
Data Scraper Agent
构建一个生产就绪的、AI驱动的数据收集代理,适用于任何公共数据源。
按计划运行,使用免费LLM丰富结果,存储到数据库,并随时间改进。
技术栈:Python · Gemini Flash(免费)· GitHub Actions(免费)· Notion / Sheets / Supabase
何时激活
- 用户希望收集或监控任何公共网站或API
- 用户说“构建一个检查...的机器人”、“帮我监控X”、“从...收集数据”
- 用户希望跟踪工作、价格、新闻、仓库、体育比分、活动、列表
- 用户询问如何在不支付托管费用的情况下自动化数据收集
- 用户希望一个基于其决策随时间变得更智能的代理
核心概念
三层架构
每个数据收集代理都有三层:
收集 → 丰富 → 存储
│ │ │
爬虫 AI (LLM) 数据库
按计划 评分/ Notion /
运行 总结 Sheets /
和分类 Supabase
免费技术栈
| 层 | 工具 | 原因 |
|---|---|---|
| 爬取 | requests + BeautifulSoup |
零成本,覆盖80%的公共网站 |
| JS渲染网站 | playwright(免费) |
当HTML获取失败时 |
| AI丰富 | Gemini Flash via REST API | 每天500次请求,100万token——免费 |
| 存储 | Notion API | 免费层,优秀的审查界面 |
| 调度 | GitHub Actions cron | 公共仓库免费 |
| 学习 | 仓库中的JSON反馈文件 | 零基础设施,持久化在git中 |
AI模型回退链
构建代理以在配额耗尽时自动回退到Gemini模型:
gemini-2.0-flash-lite (30 RPM) →
gemini-2.0-flash (15 RPM) →
gemini-2.5-flash (10 RPM) →
gemini-flash-lite-latest (回退)
批量API调用以提高效率
永远不要为每个项目调用一次LLM。始终批量处理:
# 错误:33个项目需要33次API调用
for item in items:
result = call_ai(item) # 33次调用 → 达到速率限制
# 正确:33个项目只需7次API调用(批量大小5)
for batch in chunks(items, size=5):
results = call_ai(batch) # 7次调用 → 保持在免费层内
工作流程
步骤1:理解目标
询问用户:
- 收集什么:“什么数据源?URL / API / RSS / 公共端点?”
- 提取什么:“哪些字段重要?标题、价格、URL、日期、评分?”
- 如何存储:“结果应该放在哪里?Notion、Google Sheets、Supabase还是本地文件?”
- 如何丰富:“你希望AI对每个项目进行评分、总结、分类或匹配吗?”
- 频率:“应该多久运行一次?每小时、每天、每周?”
常见示例提示:
- 招聘网站 → 根据简历评分相关性
- 产品价格 → 降价提醒
- GitHub仓库 → 总结新版本
- 新闻源 → 按主题+情感分类
- 体育结果 → 提取统计数据到追踪器
- 活动日历 → 按兴趣筛选
步骤2:设计收集架构
为用户生成以下目录结构:
my-agent/
├── config.yaml # 用户自定义(关键词、过滤器、偏好)
├── profile/
│ └── context.md # AI使用的用户上下文(简历、兴趣、标准)
├── scraper/
│ ├── __init__.py
│ ├── main.py # 编排器:爬取 → 丰富 → 存储
│ ├── filters.py # 基于规则的预过滤(快速,在AI之前)
│ └── sources/
│ ├── __init__.py
│ └── source_name.py # 每个数据源一个文件
├── ai/
│ ├── __init__.py
│ ├── client.py # Gemini REST客户端,带模型回退
│ ├── pipeline.py # 批量AI分析
│ ├── jd_fetcher.py # 从URL获取完整内容(可选)
│ └── memory.py # 从用户反馈中学习
├── storage/
│ ├── __init__.py
│ └── notion_sync.py # 或 sheets_sync.py / supabase_sync.py
├── data/
│ └── feedback.json # 用户决策历史(自动更新)
├── .env.example
├── setup.py # 一次性数据库/模式创建
├── enrich_existing.py # 对旧行回填AI评分
├── requirements.txt
└── .github/
└── workflows/
└── scraper.yml # GitHub Actions调度
步骤3:构建源连接器
任何数据源的模板:
# scraper/sources/my_source.py
"""
[源名称] — 从[哪里]收集[什么]。
方法:[REST API / HTML爬取 / RSS源]
"""
import requests
from bs4 import BeautifulSoup
from datetime import datetime, timezone
from scraper.filters import is_relevant
HEADERS = {
"User-Agent": "Mozilla/5.0 (compatible; research-bot/1.0)",
}
def fetch() -> list[dict]:
"""
返回具有一致模式的项目列表。
每个项目至少必须包含:name, url, date_found。
"""
results = []
# ---- REST API源 ----
resp = requests.get("https://api.example.com/items", headers=HEADERS, timeout=15)
if resp.status_code == 200:
for item in resp.json().get("results", []):
if not is_relevant(item.get("title", "")):
continue
results.append(_normalise(item))
return results
def _normalise(raw: dict) -> dict:
"""将原始API/HTML数据转换为标准模式。"""
return {
"name": raw.get("title", ""),
"url": raw.get("link", ""),
"source": "MySource",
"date_found": datetime.now(timezone.utc).date().isoformat(),
# 在此添加领域特定字段
}
HTML获取模式:
soup = BeautifulSoup(resp.text, "lxml")
for card in soup.select("[class*='listing']"):
title = card.select_one("h2, h3").get_text(strip=True)
link = card.select_one("a")["href"]
if not link.startswith("http"):
link = f"https://example.com{link}"
RSS源模式:
import xml.etree.ElementTree as ET
root = ET.fromstring(resp.text)
for item in root.findall(".//item"):
title = item.findtext("title", "")
link = item.findtext("link", "")
步骤4:构建Gemini AI客户端
# ai/client.py
import os, json, time, requests
_last_call = 0.0
MODEL_FALLBACK = [
"gemini-2.0-flash-lite",
"gemini-2.0-flash",
"gemini-2.5-flash",
"gemini-flash-lite-latest",
]
def generate(prompt: str, model: str = "", rate_limit: float = 7.0) -> dict:
"""调用Gemini,在429时自动回退。返回解析后的JSON或{}。"""
global _last_call
api_key = os.environ.get("GEMINI_API_KEY", "")
if not api_key:
return {}
elapsed = time.time() - _last_call
if elapsed < rate_limit:
time.sleep(rate_limit - elapsed)
models = [model] + [m for m in MODEL_FALLBACK if m != model] if model else MODEL_FALLBACK
_last_call = time.time()
for m in models:
url = f"https://generativelanguage.googleapis.com/v1beta/models/{m}:generateContent?key={api_key}"
payload = {
"contents": [{"parts": [{"text": prompt}]}],
"generationConfig": {
"responseMimeType": "application/json",
"temperature": 0.3,
"maxOutputTokens": 2048,
},
}
try:
resp = requests.post(url, json=payload, timeout=30)
if resp.status_code == 200:
return _parse(resp)
if resp.status_code in (429, 404):
time.sleep(1)
continue
return {}
except requests.RequestException:
return {}
return {}
def _parse(resp) -> dict:
try:
text = (
resp.json()
.get("candidates", [{}])[0]
.get("content", {})
.get("parts", [{}])[0]
.get("text", "")
.strip()
)
if text.startswith("```"):
text = text.split("\n", 1)[-1].rsplit("```", 1)[0]
return json.loads(text)
except (json.JSONDecodeError, KeyError):
return {}
步骤5:构建AI管道(批量)
# ai/pipeline.py
import json
import yaml
from pathlib import Path
from ai.client import generate
def analyse_batch(items: list[dict], context: str = "", preference_prompt: str = "") -> list[dict]:
"""批量分析项目。返回带有AI字段丰富后的项目。"""
config = yaml.safe_load((Path(__file__).parent.parent / "config.yaml").read_text())
model = config.get("ai", {}).get("model", "gemini-2.5-flash")
rate_limit = config.get("ai", {}).get("rate_limit_seconds", 7.0)
min_score = config.get("ai", {}).get("min_score", 0)
batch_size = config.get("ai", {}).get("batch_size", 5)
batches = [items[i:i + batch_size] for i in range(0, len(items), batch_size)]
print(f" [AI] {len(items)} 个项目 → {len(batches)} 次API调用")
enriched = []
for i, batch in enumerate(batches):
print(f" [AI] 批次 {i + 1}/{len(batches)}...")
prompt = _build_prompt(batch, context, preference_prompt, config)
result = generate(prompt, model=model, rate_limit=rate_limit)
analyses = result.get("analyses", [])
for j, item in enumerate(batch):
ai = analyses[j] if j < len(analyses) else {}
if ai:
score = max(0, min(100, int(ai.get("score", 0))))
if min_score and score < min_score:
continue
enriched.append({**item, "ai_score": score, "ai_summary": ai.get("summary", ""), "ai_notes": ai.get("notes", "")})
else:
enriched.append(item)
return enriched
def _build_prompt(batch, context, preference_prompt, config):
priorities = config.get("priorities", [])
items_text = "\n\n".join(
f"项目 {i+1}: {json.dumps({k: v for k, v in item.items() if not k.startswith('_')})}"
for i, item in enumerate(batch)
)
return f"""分析以下{len(batch)}个项目并返回一个JSON对象。
# 项目
{items_text}
# 用户上下文
{context[:800] if context else "未提供"}
# 用户优先级
{chr(10).join(f"- {p}" for p in priorities)}
{preference_prompt}
# 指令
返回:{{"analyses": [{{"score": <0-100>, "summary": "<2句话>", "notes": "<为什么匹配或不匹配>"}} for each item in order]}}
简洁。评分90+=极佳匹配,70-89=良好,50-69=一般,<50=弱。"""
步骤6:构建反馈学习系统
# ai/memory.py
"""从用户决策中学习以改进未来评分。"""
import json
from pathlib import Path
FEEDBACK_PATH = Path(__file__).parent.parent / "data" / "feedback.json"
def load_feedback() -> dict:
if FEEDBACK_PATH.exists():
try:
return json.loads(FEEDBACK_PATH.read_text())
except (json.JSONDecodeError, OSError):
pass
return {"positive": [], "negative": []}
def save_feedback(fb: dict):
FEEDBACK_PATH.parent.mkdir(parents=True, exist_ok=True)
FEEDBACK_PATH.write_text(json.dumps(fb, indent=2))
def build_preference_prompt(feedback: dict, max_examples: int = 15) -> str:
"""将反馈历史转换为提示偏差部分。"""
lines = []
if feedback.get("positive"):
lines.append("# 用户喜欢的项目(正面信号):")
for e in feedback["positive"][-max_examples:]:
lines.append(f"- {e}")
if feedback.get("negative"):
lines.append("\n# 用户跳过/拒绝的项目(负面信号):")
for e in feedback["negative"][-max_examples:]:
lines.append(f"- {e}")
if lines:
lines.append("\n使用这些模式来偏置新项目的评分。")
return "\n".join(lines)
**与存储层集成:**每次运行后,查询数据库获取正面/负面状态的项目,并使用提取的模式调用save_feedback()。
步骤7:构建存储(Notion示例)
# storage/notion_sync.py
import os
from notion_client import Client
from notion_client.errors import APIResponseError
_client = None
def get_client():
global _client
if _client is None:
_client = Client(auth=os.environ["NOTION_TOKEN"])
return _client
def get_existing_urls(db_id: str) -> set[str]:
"""获取所有已存储的URL——用于去重。"""
client, seen, cursor = get_client(), set(), None
while True:
resp = client.databases.query(database_id=db_id, page_size=100, **{"start_cursor": cursor} if cursor else {})
for page in resp["results"]:
url = page["properties"].get("URL", {}).get("url", "")
if url: seen.add(url)
if not resp["has_more"]: break
cursor = resp["next_cursor"]
return seen
def push_item(db_id: str, item: dict) -> bool:
"""将一个项目推送到Notion。成功返回True。"""
props = {
"Name": {"title": [{"text": {"content": item.get("name", "")[:100]}}]},
"URL": {"url": item.get("url")},
"Source": {"select": {"name": item.get("source", "Unknown")}},
"Date Found": {"date": {"start": item.get("date_found")}},
"Status": {"select": {"name": "New"}},
}
# AI字段
if item.get("ai_score") is not None:
props["AI Score"] = {"number": item["ai_score"]}
if item.get("ai_summary"):
props["Summary"] = {"rich_text": [{"text": {"content": item["ai_summary"][:2000]}}]}
if item.get("ai_notes"):
props["Notes"] = {"rich_text": [{"text": {"content": item["ai_notes"][:2000]}}]}
try:
get_client().pages.create(parent={"database_id": db_id}, properties=props)
return True
except APIResponseError as e:
print(f"[notion] 推送失败:{e}")
return False
def sync(db_id: str, items: list[dict]) -> tuple[int, int]:
existing = get_existing_urls(db_id)
added = skipped = 0
for item in items:
if item.get("url") in existing:
skipped += 1; continue
if push_item(db_id, item):
added += 1; existing.add(item["url"])
else:
skipped += 1
return added, skipped
步骤8:在main.py中编排
# scraper/main.py
import os, sys, yaml
from pathlib import Path
from dotenv import load_dotenv
load_dotenv()
from scraper.sources import my_source # 添加你的源
# 注意:此示例使用Notion。如果storage.provider是"sheets"或"supabase",
# 请将此导入替换为storage.sheets_sync或storage.supabase_sync,并更新
# 环境变量和sync()调用。
from storage.notion_sync import sync
SOURCES = [
("My Source", my_source.fetch),
]
def ai_enabled():
return bool(os.environ.get("GEMINI_API_KEY"))
def main():
config = yaml.safe_load((Path(__file__).parent.parent / "config.yaml").read_text())
provider = config.get("storage", {}).get("provider", "notion")
# 根据提供者从环境变量解析存储目标标识符
if provider == "notion":
db_id = os.environ.get("NOTION_DATABASE_ID")
if not db_id:
print("错误:NOTION_DATABASE_ID未设置"); sys.exit(1)
else:
# 在此扩展以支持sheets(SHEET_ID)或supabase(SUPABASE_TABLE)等。
print(f"错误:提供者'{provider}'尚未在main.py中连接"); sys.exit(1)
config = yaml.safe_load((Path(__file__).parent.parent / "config.yaml").read_text())
all_items = []
for name, fetch_fn in SOURCES:
try:
items = fetch_fn()
print(f"[{name}] {len(items)} 个项目")
all_items.extend(items)
except Exception as e:
print(f"[{name}] 失败:{e}")
# 按URL去重
seen, deduped = set(), []
for item in all_items:
if (url := item.get("url", "")) and url not in seen:
seen.add(url); deduped.append(item)
print(f"唯一项目:{len(deduped)}")
if ai_enabled() and deduped:
from ai.memory import load_feedback, build_preference_prompt
from ai.pipeline import analyse_batch
# load_feedback()读取由反馈同步脚本写入的data/feedback.json。
# 要保持最新,请实现一个单独的feedback_sync.py,查询存储提供者
# 获取正面/负面状态的项目,并调用save_feedback()。
feedback = load_feedback()
preference = build_preference_prompt(feedback)
context_path = Path(__file__).parent.parent / "profile" / "context.md"
context = context_path.read_text() if context_path.exists() else ""
deduped = analyse_batch(deduped, context=context, preference_prompt=preference)
else:
print("[AI] 跳过——GEMINI_API_KEY未设置")
added, skipped = sync(db_id, deduped)
print(f"完成——新增{added}个,已存在{skipped}个")
if __name__ == "__main__":
main()
步骤9:GitHub Actions工作流
# .github/workflows/scraper.yml
name: Data Scraper Agent
on:
schedule:
- cron: "0 */3 * * *" # 每3小时——根据需要进行调整
workflow_dispatch: # 允许手动触发
permissions:
contents: write # 反馈历史提交步骤所需
jobs:
scrape:
runs-on: ubuntu-latest
timeout-minutes: 20
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.11"
cache: "pip"
- run: pip install -r requirements.txt
# 如果requirements.txt中启用了Playwright,请取消注释
# - name: Install Playwright browsers
# run: python -m playwright install chromium --with-deps
- name: Run agent
env:
NOTION_TOKEN: ${{ secrets.NOTION_TOKEN }}
NOTION_DATABASE_ID: ${{ secrets.NOTION_DATABASE_ID }}
GEMINI_API_KEY: ${{ secrets.GEMINI_API_KEY }}
run: python -m scraper.main
- name: Commit feedback history
run: |
git config user.name "github-actions[bot]"
git config user.email "github-actions[bot]@users.noreply.github.com"
git add data/feedback.json || true
git diff --cached --quiet || git commit -m "chore: update feedback history"
git push
步骤10:config.yaml模板
# 自定义此文件——无需修改代码
# 收集什么(AI之前的预过滤)
filters:
required_keywords: [] # 项目必须至少包含一个
blocked_keywords: [] # 项目不得包含任何
# 你的优先级——AI用于评分
priorities:
- "示例优先级1"
- "示例优先级2"
# 存储
storage:
provider: "notion" # notion | sheets | supabase | sqlite
# 反馈学习
feedback:
positive_statuses: ["Saved", "Applied", "Interested"]
negative_statuses: ["Skip", "Rejected", "Not relevant"]
# AI设置
ai:
enabled: true
model: "gemini-2.5-flash"
min_score: 0 # 过滤掉低于此评分的项目
rate_limit_seconds: 7 # API调用之间的秒数
batch_size: 5 # 每次API调用的项目数
常见爬取模式
模式1:REST API(最简单)
resp = requests.get(url, params={"q": query}, headers=HEADERS, timeout=15)
items = resp.json().get("results", [])
模式2:HTML爬取
soup = BeautifulSoup(resp.text, "lxml")
for card in soup.select(".listing-card"):
title = card.select_one("h2").get_text(strip=True)
href = card.select_one("a")["href"]
模式3:RSS源
import xml.etree.ElementTree as ET
root = ET.fromstring(resp.text)
for item in root.findall(".//item"):
title = item.findtext("title", "")
link = item.findtext("link", "")
pub_date = item.findtext("pubDate", "")
模式4:分页API
page = 1
while True:
resp = requests.get(url, params={"page": page, "limit": 50}, timeout=15)
data = resp.json()
items = data.get("results", [])
if not items:
break
for item in items:
results.append(_normalise(item))
if not data.get("has_more"):
break
page += 1
模式5:JS渲染页面(Playwright)
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.goto(url)
page.wait_for_selector(".listing")
html = page.content()
browser.close()
soup = BeautifulSoup(html, "lxml")
应避免的反模式
| 反模式 | 问题 | 修复 |
|---|---|---|
| 每个项目一次LLM调用 | 立即达到速率限制 | 每次调用批量处理5个项目 |
| 代码中硬编码关键词 | 不可重用 | 将所有配置移至config.yaml |
| 无速率限制的爬取 | IP被封 | 在请求之间添加time.sleep(1) |
| 在代码中存储密钥 | 安全风险 | 始终使用.env + GitHub Secrets |
| 无去重 | 重复行堆积 | 在推送前始终检查URL |
忽略robots.txt |
法律/道德风险 | 遵守爬取规则;尽可能使用公共API |
使用requests处理JS渲染网站 |
空响应 | 使用Playwright或查找底层API |
maxOutputTokens太低 |
JSON截断,解析错误 | 批量响应使用2048+ |
免费层限制参考
| 服务 | 免费限制 | 典型使用 |
|---|---|---|
| Gemini Flash Lite | 30 RPM, 1500 RPD | 3小时间隔约56次请求/天 |
| Gemini 2.0 Flash | 15 RPM, 1500 RPD | 良好的回退 |
| Gemini 2.5 Flash | 10 RPM, 500 RPD | 谨慎使用 |
| GitHub Actions | 无限(公共仓库) | 约20分钟/天 |
| Notion API | 无限 | 约200次写入/天 |
| Supabase | 500MB数据库,2GB传输 | 适用于大多数代理 |
| Google Sheets API | 300次请求/分钟 | 适用于小型代理 |
依赖模板
requests==2.31.0
beautifulsoup4==4.12.3
lxml==5.1.0
python-dotenv==1.0.1
pyyaml==6.0.2
notion-client==2.2.1 # 如果使用Notion
# playwright==1.40.0 # 如果使用JS渲染网站,取消注释
质量检查清单
在标记代理完成之前:
- [ ]
config.yaml控制所有面向用户的设置——无硬编码值 - [ ]
profile/context.md保存用户特定上下文,用于AI匹配 - [ ] 每次存储推送前按URL去重
- [ ] Gemini客户端具有模型回退链(4个模型)
- [ ] 每次API调用批量大小≤5个项目
- [ ]
maxOutputTokens≥ 2048 - [ ]
.env在.gitignore中 - [ ] 提供
.env.example用于入门 - [ ]
setup.py在首次运行时创建数据库模式 - [ ]
enrich_existing.py回填旧行的AI评分 - [ ] GitHub Actions工作流在每次运行后提交
feedback.json - [ ] README涵盖:5分钟内设置、所需密钥、自定义
实际示例
"为我构建一个代理,监控Hacker News上AI创业融资新闻"
"从3个电商网站抓取产品价格,并在降价时提醒"
"跟踪标记为'llm'或'agents'的新GitHub仓库——总结每个仓库"
"从LinkedIn和Cutshort收集Chief of Staff职位列表到Notion"
"监控一个子版块中提及我公司的帖子——分类情感"
"每天从arXiv抓取我关心的主题的新学术论文"
"跟踪体育赛事结果并在Google Sheets中维护运行表"
"构建一个房地产列表观察器——在低于₹1 Cr的新房产上提醒"
参考实现
一个使用此确切架构构建的完整工作代理将从4个以上来源收集数据,批量调用Gemini,从存储在Notion中的已申请/已拒绝决策中学习,并在GitHub Actions上100%免费运行。按照上述步骤1-9构建你自己的代理。






