A beginner-friendly guide for Linux. No programming knowledge needed - just follow each step in order. Written for a fresh server (or a normal desktop) with nothing installed.
Read this first. This guide assumes your computer has nothing installed for this project - no unzip, no curl, no uv, and no browser. Every tool gets its own install step below. Do the steps in order and do not skip any. The five most common mistakes are:
cd tiktok fails - because the zip has no tiktok folder inside it. Step 3 fixes this.
uv sync fails - because uv is not installed by default. Step 4 fixes this.
Trying to install Python: you don't need to. uv supplies Python itself. The python3.12 package does not exist on many Ubuntu/Debian versions.
Skipping xvfb: you must install it. Without xvfb-run headless fetching is impossible. Step 5 fixes this.
Running headless first in Step 8 - wrong order. You must run Step 8A headed-under-Xvfb first (it will fail with 0 videos on purpose — ~90% of the time), and only then run Step 8B headless (retry twice for ~11/12 videos).
A small program that:
Opens a real browser (Brave or Chrome) and visits a TikTok search results page
Scrolls down so more videos load
Saves the page HTML locally
Pulls out each video (URL, author, description, likes, hashtags, thumbnail) and prints it
TikTok builds its search page with JavaScript, so the raw page has no results in it. We use a real browser controlled through Chrome DevTools Protocol (CDP). The browser loads the page, scrolls, and grabs the finished HTML. After that first fetch, everything runs offline against the saved file.
The browser runs headless (invisible) for the final fetch, but headless cannot work on its own. You must first do one headed init run under a virtual screen called Xvfb (Step 5 + Step 8A). That headed run invokes TikTok's JavaScript to verify your session and cookies — it fails with 0 videos ~90% of the time on purpose. Only after it fails do you run the headless fetch (Step 8B), which succeeds on the second try with about 11/12 videos.
You should have received one file: tiktok.zip. It already contains the whole program - including the parser, the API, and the command-line tool. You do not copy code from any other project and you do not type the code by hand.
A terminal is a window where you type commands instead of clicking buttons. Every command in this guide is typed here, then you press Enter.
If you are on a server (a droplet/VPS), you are already inside a terminal after you connect. Skip to the check below.
To open it on a Linux desktop:
Press Ctrl + Alt + T at the same time, or
Open your applications menu and search for Terminal, Console, or GNOME Terminal
Check it works: type this and press Enter:
pwdIt prints the folder you are currently in. If it prints a path like /home/yourname, the terminal works. On a server, echo $DISPLAY printing nothing is normal - the default headless browser does not need a screen.
Two words you will see a lot:
sudo - "run this as administrator." It will ask for your login password. Nothing appears on screen while you type the password - that is normal. Type it and press Enter.
apt - Ubuntu/Debian's built-in installer for system tools. If you use Fedora, replace apt install with dnf install.
unzip opens your tiktok.zip file. curl downloads the uv installer. Neither is installed by default on a fresh Linux system, so install them now.
sudo apt update
sudo apt install -y unzip curlFedora users:
sudo dnf install -y unzip curlVerify:
unzip -v
curl --versionEach should print version information. If either says command not found, run the install command again and read the error message.
tiktok)cd tiktok problemThe zip was built with its files at the top level - meaning there is no tiktok folder inside the zip. If you just double-click "Extract Here," the files (main.py, api.py, ...) get dumped loose into the current folder, and later cd tiktok fails with No such file or directory.
The fix: create a folder named tiktok first, then tell unzip to put the files inside it.
# 1. Go to where the uploaded zip lives (usually your home folder)
cd ~
# 2. Create the destination folder
mkdir -p tiktok
# 3. Extract the zip's contents INTO that folder (-d means destination)
unzip -o tiktok.zip -d tiktok
# 4. Enter the folder
cd tiktok
# 5. Confirm the files are here
lsAfter step 5, ls must show files such as main.py, pyproject.toml, tiktok_parse.py, and folders browser, models, data.
Already extracted the files loose somewhere? Move them into a proper folder:
mkdir -p ~/tiktok
mv api.py constants.py fetch_sample.py main.py pyproject.toml \
tiktok_parse.py uv.lock sample.html index.html TUTORIAL.md \
browser models data ~/tiktok/
cd ~/tiktokErrors about files that are "not found" are fine if those files are already in place.
If unzip says cannot find or open tiktok.zip: the zip is not in your home folder. Find it with find ~ -name "tiktok.zip", then cd into the folder it printed before running unzip again.
From now on, every command is run while inside ~/tiktok. If you ever get lost, run cd ~/tiktok to return. Confirm with pwd - it should end in /tiktok.
uv sync problem (and means you never install Python by hand)uv is a fast tool that installs the Python libraries this project needs. It also provides Python itself, so you do not install Python - and you should not try to. The python3.12 apt package does not exist on Ubuntu 22.04 / 26.04 or Debian 12 / 13, so trying to install it just fails. uv avoids all of that.
curl -LsSf https://astral.sh/uv/install.sh | shThis uses curl (installed in Step 2) to download the installer and run it. It puts uv in the folder ~/.local/bin.
Now close and reopen your terminal so it can find the new program. (Or run source $HOME/.local/bin/env.)
Verify:
uv --versionShould show something like uv 0.x.x.
If uv: command not found still appears, add the folder to your PATH, then reload it:
echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.bashrc
source ~/.bashrc
uv --versionThe program controls a real browser through a debugging port. Either browser works.
Tested on Brave. This scraper was built and verified with Brave. It is the recommended browser and the default path (/usr/bin/brave-browser) already points at it.
Brave (recommended) - official repository:
sudo apt update
sudo apt install -y curl gpg
sudo curl -fsSLo /usr/share/keyrings/brave-browser-archive-keyring.gpg \
https://brave-browser-apt-release.s3.brave.com/brave-browser-archive-keyring.gpg
echo "deb [signed-by=/usr/share/keyrings/brave-browser-archive-keyring.gpg] \
https://brave-browser-apt-release.s3.brave.com/ stable main" \
| sudo tee /etc/apt/sources.list.d/brave-browser-release.list
sudo apt update
sudo apt install -y brave-browserMore detail is available at brave.com/linux.
Google Chrome instead:
sudo apt install -y wget
wget https://dl.google.com/linux/direct/google-chrome-stable_current_amd64.deb
sudo apt install -y ./google-chrome-stable_current_amd64.debInstalling this .deb is what creates /usr/bin/google-chrome. If Chrome is installed some other way, that path may not exist.
Find the browser's location. The default in constants.py is /usr/bin/brave-browser. Check the real path with:
which brave-browser
which google-chromeIf it prints a different path, you will edit one line of constants.py later (the file is in the reference section).
MUST-DO - required for Step 8. Headless fetching is impossible without this. Even though the final fetch is headless, the mandatory first run (Step 8A) is headed and needs a virtual screen on a server. Install it now:
sudo apt install -y xvfb
which xvfb-run
xvfb-run --helpwhich xvfb-run must print a path (usually /usr/bin/xvfb-run). Fedora users: sudo dnf install -y xorg-x11-server-Xvfb. Step 8A will not run without this.
uv sync from inside the projectMake sure you are in the project folder first, then sync:
cd ~/tiktok
uv syncThis reads pyproject.toml and installs beautifulsoup4, lxml, httpx, pydantic, psutil and websocket-client, plus developer tools (pytest, ruff, pyright).
If uv sync fails with uv: command not found: go back to Step 4. uv sync cannot work until uv --version works.
If uv sync fails with network/download errors: check your internet connection and run uv sync again. The first run downloads everything.
Linux only - pyright dependency: If uv run pyright fails with libatomic.so.1: cannot open shared object file:
sudo apt install -y libatomic1uv run python main.pyWhat happens:
A sample.html ships with the project, so it is parsed directly and no browser opens
The parser reads the HTML and prints each video
Runs after that: the browser never opens unless you ask with --fetch. The program reads sample.html and parses it instantly.
Using cached /home/<you>/tiktok/sample.html (533498 bytes)
============================================================
SEARCH: 'new trend' (12 videos)
============================================================
1. @eyochico (grimes)
url : https://www.tiktok.com/@eyochico/video/7585351064929553677
video id : 7585351064929553677
desc : holy duo #foryou #adrianalima #meganfox #mogging #targetaudience
likes : 237800 comments: - shares: - plays: -
posted : 2025-12-19
hashtags : #foryou #adrianalima #meganfox #mogging #targetaudience
cover : https://p16-common-sign.tiktokcdn.com/tos-useast5-p-0068-tx/...
============================================================
Done.The number of videos changes from one capture to the next. The sample bundled in this zip holds 12 videos. A live fetch done as 8A headed-init (must fail) → 8B headless retry returns at most 11/12 videos. The SEARCH: line always prints the real number. The first item is the same (@eyochico) across samples of this query.
# A different query always needs --fetch — redo 8A then 8B:
xvfb-run -a uv run python main.py --query "dance challenge" --fetch --headed
# expect 0 videos (init), then:
uv run python main.py --query "dance challenge" --fetch
# if 0 videos, run the same headless command once more (max ~11/12)If you run --query "something" without --fetch, the program stops and tells you there is no downloaded cache for it. That is on purpose - it will not show you the wrong query's data.
Order is mandatory. Do not reverse it. Steps 1-7 were normal setup. Now: 8A headed-under-Xvfb first (it MUST fail), then 8B headless (retry for ~11/12 videos). Running headless before 8A will not work.
xvfb-run -a uv run python main.py --fetch --headedThis step does not give you videos. Its whole purpose is to invoke TikTok's JavaScript to verify your session and cookies and initialize the page/profile. Expect:
All scrolls show 0 video links rendered
Program stops with Fetch returned 0 videos
Failed page goes to sample.failed.html — your good sample.html is not overwritten
If 8A shows 0 videos, it worked as intended. Move to 8B. This requires xvfb from Step 5 — without xvfb-run this step (and therefore headless) is impossible.
# Run only after 8A has failed once:
uv run python main.py --fetchA Brave/Chrome window stays invisible (headless) - that is normal. It loads the page, waits about 12 seconds, scrolls, saves sample.html, then closes.
8B still fails the first time — this is also expected. If the first 8B run returns 0 videos, run the exact same headless command a second time:
uv run python main.py --fetchThe second try generates at most 11/12 videos. That is success. Do not expect 47-83. If it is still empty after two 8B tries, see Troubleshooting.
You will see scroll messages like:
scroll 1/6: ~12 video links rendered
scroll 2/6: ~31 video links renderedIn 8A every scroll showing 0 is expected (init failure). In 8B retry the first scroll already shows a non-zero number (about 11/12 headless). If 8B still shows 0 on every scroll twice in a row, you got a login wall/captcha — see Troubleshooting.
If you also have the Instagram project, you can share its browser profile so both projects use the same cookies/session. Skip this whole step if you do not have the Instagram project - the browser creates its own profile on first run.
# run from inside ~/tiktok
cd ~/tiktok
mkdir -p ../instagram/browser/browser_profile
ln -s ../../instagram/browser/browser_profile browser/browser_profile
ls -l browser/browser_profile # verify the arrow points to a real folderIf the link already exists (or is broken) you get a FileExistsError. Remove it and create it again:
rm browser/browser_profile
ln -s ../../instagram/browser/browser_profile browser/browser_profilecd tiktok - "No such file or directory"
The zip has no tiktok folder inside it. Redo Step 3: mkdir -p tiktok, then unzip -o tiktok.zip -d tiktok, then cd tiktok. Run ls to see where you are.
"Command not found: uv" or uv sync fails
You skipped Step 4. Install uv with the curl command, close and reopen the terminal, and check uv --version. If it still fails, add ~/.local/bin to your PATH (Step 4).
"Command not found: python3.12"
You do not need system Python. uv supplies Python itself. Do not run apt install python3.12 - that package does not exist on Ubuntu 22.04 / 26.04 or Debian 12 / 13.
"Command not found: curl"
Install it: sudo apt install -y curl (Step 2).
"Command not found: unzip"
Install it: sudo apt install -y unzip (Step 2).
Browser doesn't open / "connection refused" / hangs on "Waiting for browser to be ready"
The program now times out instead of hanging forever and prints the browser log. Check constants.py: find your browser path with which brave-browser or which google-chrome and set DEBUG_BROWSER_PATH to match.
Error mentions Missing X server or $DISPLAY
You ran Step 8A without Xvfb. Install it: sudo apt install -y xvfb (Step 5) and always prefix headed with xvfb-run -a: xvfb-run -a uv run python main.py --fetch --headed. Verify with which xvfb-run.
Chrome doesn't work
The scraper was tested on Brave. Install Brave (Step 5) and set DEBUG_BROWSER_PATH = "/usr/bin/brave-browser" in constants.py. Note: /usr/bin/google-chrome only exists if you installed Chrome from the .deb.
0 videos on Step 8A (--headed under xvfb-run)
This MUST happen (~90%). 8A is init-only: it invokes TikTok JavaScript to verify session/cookies. Do not retry 8A — move to Step 8B headless: uv run python main.py --fetch.
0 videos on the first 8B (headless) run
Also expected. Run the exact same headless command a second time: uv run python main.py --fetch. The second try generates at most 11/12 videos.
0 videos even after two 8B runs
Look at sample.failed.html (the saved error page). It may be a login wall or captcha, or your network/profile is blocked. Your sample.html is safe; restore the previous one with cp sample.html.bak sample.html if needed.
Only 11/12 videos (or fewer)
That is success — headless caps at ~11/12 after the 8A init + 8B retry sequence. Do not switch back to --headed for more; if you get fewer than 11, increase SCROLL_STEPS and SCROLL_PAUSE in constants.py, then redo 8A-fail → 8B-retry with --fetch.
Port 9999 already in use
Another browser is still open. Close it and try again, or pass a different debug_port to TikTokClient in your own script (the port is honored, not hard-coded).
FileExistsError on browser_profile
The symlink is broken. Remove it and recreate it (Step 9).
"No module named 'models'"
Make sure you ran unzip -o tiktok.zip -d tiktok fully, so the models/ and browser/ folders exist. Confirm with ls.
pytest says "no tests ran"
Expected - the zip ships no tests/ folder. Not a failure.
Everything below already ships inside tiktok.zip - you do not need to create, copy, or paste any of it. This section is here only so you know what each file does and can see the code if you are curious. None of it comes from the Instagram project.
pyproject.toml[project]
name = "tiktok-scraper"
version = "0.1.0"
description = "TikTok search result scraper using Brave CDP browser"
readme = "README.md"
requires-python = ">=3.12"
dependencies = [
"beautifulsoup4>=4.15.0",
"httpx>=0.28.1",
"lxml>=6.1.1",
"psutil>=7.0.0",
"pydantic>=2.13.4",
"websocket-client>=1.9.0",
]
[dependency-groups]
dev = [
"pyright>=1.1.410",
"pytest>=9.0.3",
"ruff>=0.15.16",
]
[tool.ruff]
target-version = "py312"
line-length = 88
[tool.ruff.lint]
fixable = ["ALL"]
select = ["I", "B", "E"]
[tool.pyright]
typeCheckingMode = "strict"
venvPath = "."
venv = ".venv"
pythonVersion = "3.12"
[tool.pytest.ini_options]
testpaths = ["tests"]
pythonpath = ["."]constants.pyfrom pathlib import Path
# Paths
HOME = Path.home()
BASE_DIR = Path(__file__).parent
# Browser
DEBUG_BROWSER_PATH = "/usr/bin/brave-browser"
# Symlinked to ../../instagram/browser/browser_profile so the TikTok runs reuse
# the same browser session / cookies as the Instagram project.
USER_PROFILE_DIR = BASE_DIR / "browser" / "browser_profile"
# Browser Configs
WIN_W = 720
WIN_H = 760
# TikTok
TIKTOK_BASE = "https://www.tiktok.com"
DEFAULT_QUERY = "new trend"
SAMPLE_HTML_PATH = BASE_DIR / "sample.html"
# Kept so a failed live fetch never destroys a good sample.html.
SAMPLE_HTML_BACKUP_PATH = BASE_DIR / "sample.html.bak"
# Remembers which query produced sample.html, so a mismatched --query is caught.
SAMPLE_META_PATH = BASE_DIR / "sample.meta.json"
# Browser stdout/stderr are captured here instead of being thrown away.
BROWSER_LOG_PATH = BASE_DIR / "browser" / "browser.log"
# Seconds to wait for the browser debug port before giving up.
BROWSER_READY_TIMEOUT = 30
# How long to let the client-rendered search page settle before grabbing HTML.
# Note: headless mode may render fewer results than headed mode. Use --headed
# (under xvfb-run on a server) when you need the full result set.
RENDER_TIMEOUT = 12
# Scroll settings used to trigger lazy-loaded search results.
SCROLL_STEPS = 6
SCROLL_DELTA_Y = 4000
SCROLL_PAUSE = 2.0If your browser is not at /usr/bin/brave-browser, edit the DEBUG_BROWSER_PATH line. Find the path with which brave-browser or which google-chrome.
models/search.py - the data shapefrom pydantic import BaseModel, Field
class TikTokVideoItem(BaseModel):
"""One video item parsed from a TikTok search result page.
All fields are optional / defaulted so a card can still be captured even
when the rendered DOM omits a specific value.
Note: TikTok renders the engagement number inside the search card with a
heart (like) icon but labels the element ``data-e2e="video-views"``. It is
mapped to :attr:`likes`. Comments / shares / plays are not present in the
search-result DOM, so they stay ``None`` unless a future layout adds them.
"""
video_id: str | None = None
video_url: str | None = None
description: str | None = None
hashtags: list[str] = Field(default_factory=list)
author_username: str | None = None
author_nickname: str | None = None
author_url: str | None = None
author_avatar_url: str | None = None
author_verified: bool | None = None
likes: int | None = None
comments: int | None = None
shares: int | None = None
plays: int | None = None
cover_url: str | None = None
duration: int | None = None
create_time: int | None = None
create_time_text: str | None = None
music_title: str | None = None
top_liked: bool = False
likes_raw: str | None = None
comments_raw: str | None = None
shares_raw: str | None = None
plays_raw: str | None = None
class TikTokSearchData(BaseModel):
query: str | None = None
result_count: int = 0
videos: list[TikTokVideoItem] = Field(default_factory=list[TikTokVideoItem])tiktok_parse.py - how the page is readThis finds the result list and reads each card. Its core selectors:
| Selector | What it is |
|---|---|
| [data-e2e="search_top-item-list"] | the results container |
| [data-e2e="search_top-item"] | one video card (link + likes) |
| [data-e2e="search-card-desc"] | caption + author block |
| [data-e2e="search-card-video-caption"] | caption text + hashtags |
| [data-e2e="search-card-user-link"] | author link (/@username) |
| [data-e2e="search-card-user-unique-id"] | author nickname |
| [data-e2e="video-views"] | the likes number (next to a heart icon) |
Counts are abbreviated. TikTok writes 39.4K or 1.2M. The parser converts them to 39400 and 1200000 so you can sort and add them up.
Full source:
"""
Parse a rendered TikTok search result page into structured data.
Extraction strategy (matching the rendered DOM served today):
- Each result lives in a direct child of the ``[data-e2e="search_top-item-list"]``
container.
- The video link/cover/likes come from the ``[data-e2e="search_top-item"]`` card.
- The caption, hashtags and author come from the sibling
``[data-e2e="search-card-desc"]`` block.
TikTok search pages are client-side rendered: the server HTML is only a shell.
This parser therefore expects the *browser-rendered* DOM saved to sample.html by
``fetch_sample.py``.
Run directly to validate against the local sample.html:
uv run python tiktok_parse.py
"""
import json
import re
from bs4 import BeautifulSoup
from bs4.element import Tag
from models.search import TikTokSearchData, TikTokVideoItem
TIKTOK_BASE = "https://www.tiktok.com"
LIST_SELECTOR = '[data-e2e="search_top-item-list"]'
CARD_SELECTOR = '[data-e2e="search_top-item"]'
CAPTION_SELECTOR = '[data-e2e="search-card-video-caption"]'
USER_LINK_SELECTOR = '[data-e2e="search-card-user-link"]'
NICKNAME_SELECTOR = '[data-e2e="search-card-user-unique-id"]'
LIKE_SELECTOR = '[data-e2e="video-views"]'
HASHTAG_SELECTOR = '[data-e2e="search-common-link"]'
_VIDEO_ID_RE = re.compile(r"/video/(\d+)")
_COUNT_RE = re.compile(r"([\d.]+)\s*([KMB]?)", re.IGNORECASE)
_COUNT_MULT = {"": 1, "K": 1_000, "M": 1_000_000, "B": 1_000_000_000}
def parse_count(text: str | None) -> int | None:
"""Convert TikTok's abbreviated counts ("1.2M", "39.4K", "1760") to int."""
if not text:
return None
cleaned = text.replace(",", "").replace(" ", "").strip()
match = _COUNT_RE.search(cleaned)
if not match:
return None
try:
value = float(match.group(1))
except ValueError:
return None
multiplier = _COUNT_MULT.get(match.group(2).upper(), 1)
return int(value * multiplier)
def _select_one(node: Tag, selector: str) -> Tag | None:
return node.select_one(selector)
def _select(node: Tag, selector: str) -> list[Tag]:
return node.select(selector)
def _attr(node: Tag, name: str) -> str | None:
value = node.get(name)
return value if isinstance(value, str) else None
def _text(node: Tag | None) -> str | None:
if node is None:
return None
value = node.get_text(" ", strip=True)
return value or None
def _parse_hashtags(caption: Tag | None) -> list[str]:
if caption is None:
return []
tags: list[str] = []
for link in _select(caption, HASHTAG_SELECTOR):
href = _attr(link, "href") or ""
if href.startswith("/tag/"):
tag = link.get_text(strip=True)
if tag and tag not in tags:
tags.append(tag)
return tags
def _parse_time_text(user_link: Tag | None) -> str | None:
"""The relative/absolute post date sits next to the author link."""
if user_link is None:
return None
row = user_link.parent
if not isinstance(row, Tag):
return None
for sibling in row.find_all(recursive=False):
if sibling is user_link:
continue
text = sibling.get_text(strip=True)
if text:
return text
return None
def _parse_video_link(card: Tag) -> Tag | None:
for anchor in card.find_all("a"):
href = _attr(anchor, "href") or ""
if _VIDEO_ID_RE.search(href):
return anchor
return None
def parse_video_item(container: Tag) -> TikTokVideoItem | None:
"""Parse one search result container into a TikTokVideoItem."""
card = _select_one(container, CARD_SELECTOR)
if card is None:
return None
link = _parse_video_link(card)
if link is None:
return None
video_url = _attr(link, "href")
video_match = _VIDEO_ID_RE.search(video_url or "")
video_id = video_match.group(1) if video_match else None
cover = _select_one(link, "img")
cover_url = _attr(cover, "src") if cover is not None else None
likes_raw = _text(_select_one(card, LIKE_SELECTOR))
caption = _select_one(container, CAPTION_SELECTOR)
description = _text(caption)
hashtags = _parse_hashtags(caption)
user_link = _select_one(container, USER_LINK_SELECTOR)
author_username = None
author_url = None
author_avatar_url = None
if user_link is not None:
href = _attr(user_link, "href") or ""
if href.startswith("/@"):
author_username = href[2:]
author_url = f"{TIKTOK_BASE}{href}"
avatar = _select_one(user_link, "img")
author_avatar_url = _attr(avatar, "src") if avatar is not None else None
author_nickname = _text(_select_one(container, NICKNAME_SELECTOR))
create_time_text = _parse_time_text(user_link)
top_liked = "Top liked" in card.get_text()
return TikTokVideoItem(
video_id=video_id,
video_url=video_url,
description=description,
hashtags=hashtags,
author_username=author_username,
author_nickname=author_nickname,
author_url=author_url,
author_avatar_url=author_avatar_url,
likes=parse_count(likes_raw),
likes_raw=likes_raw,
cover_url=cover_url,
create_time_text=create_time_text,
top_liked=top_liked,
)
def _extract_query(soup: BeautifulSoup) -> str | None:
box = soup.select_one('[data-e2e="search-user-input"]')
if box is not None:
value = box.get("value")
if isinstance(value, str) and value:
return value
return None
def parse_search_page(html: str, query: str | None = None) -> TikTokSearchData:
"""Parse rendered TikTok search HTML and return structured data."""
soup = BeautifulSoup(html, "lxml")
container = soup.select_one(LIST_SELECTOR)
items: list[TikTokVideoItem] = []
seen_ids: set[str] = set()
if container is not None:
for child in container.find_all(recursive=False):
item = parse_video_item(child)
if item is None or item.video_id is None:
continue
if item.video_id in seen_ids:
continue
seen_ids.add(item.video_id)
items.append(item)
return TikTokSearchData(
query=query or _extract_query(soup),
result_count=len(items),
videos=items,
)
if __name__ == "__main__":
from constants import SAMPLE_HTML_PATH
if not SAMPLE_HTML_PATH.exists():
print("sample.html not found. Run fetch_sample.py first.")
raise SystemExit(1)
data = parse_search_page(SAMPLE_HTML_PATH.read_text(encoding="utf-8"))
print("=== Search ===")
print(f" query : {data.query}")
print(f" results : {data.result_count}")
print()
for i, video in enumerate(data.videos, 1):
print(f" {i:>2}. @{video.author_username} ({video.author_nickname})")
print(f" url : {video.video_url}")
print(f" desc : {(video.description or '')[:80]}")
print(f" likes : {video.likes} (raw {video.likes_raw!r})")
print(f" posted : {video.create_time_text}")
print(f" hashtags : {video.hashtags}")
print(f" cover : {(video.cover_url or '')[:60]}...")
print()
print(json.dumps(data.model_dump(mode="json"), ensure_ascii=False)[:500] + " ...")api.py and main.py - fetching and the command lineapi.py drives the browser (or reads sample.html) and calls the parser. main.py is the command-line tool with --query, --fetch, and --headed. Full source:
"""
TikTok search scraper client.
TikTok search pages are client-side rendered, so live fetches go through the
Brave CDP browser and capture the rendered DOM after scrolling to trigger
lazy-loaded results.
If sample.html already exists, the browser is never spawned - all parse work
runs against the local file.
The browser runs **headless** by default. On a fresh browser profile the very
first live fetch returns 0 videos (TikTok serves an error page instead of a
login wall or captcha). In that case the fetch is stopped with
:class:`FirstRunWarmupError`, which tells the caller to rerun the same command
once. The first run only warms cookies/caches.
"""
import json
import time
from pathlib import Path
from urllib.parse import quote
from browser.actions import (
browser_open_url,
dom_element_html,
dom_js_execute,
dom_scroll_at,
)
from browser.browser import get_browser_instance, kill_browser
from constants import (
RENDER_TIMEOUT,
SAMPLE_HTML_BACKUP_PATH,
SAMPLE_META_PATH,
SCROLL_DELTA_Y,
SCROLL_PAUSE,
SCROLL_STEPS,
TIKTOK_BASE,
WIN_H,
WIN_W,
)
from models.search import TikTokSearchData
from tiktok_parse import parse_search_page
VIDEO_LINK_COUNT_JS = 'document.querySelectorAll(\'a[href*="/video/"]\').length'
FIRST_RUN_MESSAGE = (
"Fetch returned 0 videos (0/0). The first fetch on a fresh browser profile "
"only warms cookies/caches and is expected to fail on this same system. "
"Rerun the exact same command once - it should then return results, e.g.:\n"
" uv run python main.py --fetch\n"
"Do not change code or config. A second run normally returns videos; if it "
"still returns 0, the profile/network is the problem, not the setup."
)
class FirstRunWarmupError(RuntimeError):
"""Raised when a live fetch returns zero videos on a cold profile."""
class TikTokClient:
def __init__(self, debug_port: int = 9999, headless: bool = True) -> None:
self.debug_port = debug_port
self.headless = headless
self._browser = None
def _ensure_browser(self) -> None:
if self._browser is None:
mode = "headless" if self.headless else "headed"
print(f"Spawning Brave browser ({mode}) ...")
self._browser = get_browser_instance(
self.debug_port, headless=self.headless
)
def _scroll_to_load(self) -> None:
assert self._browser is not None
cx, cy = WIN_W // 2, WIN_H // 2
for i in range(SCROLL_STEPS):
dom_scroll_at(self._browser, cx, cy, delta_y=SCROLL_DELTA_Y)
time.sleep(SCROLL_PAUSE)
count = dom_js_execute(self._browser, VIDEO_LINK_COUNT_JS).get("value", 0)
print(f" scroll {i + 1}/{SCROLL_STEPS}: ~{count} video links rendered")
@staticmethod
def search_url(query: str) -> str:
return f"{TIKTOK_BASE}/search?q={quote(query)}"
def fetch_search_html(self, query: str) -> str:
"""Navigate straight to the search page and return the rendered DOM.
There is deliberately no visit to tiktok.com first and no in-process
retry: a cold-profile failure is reported to the caller so the operator
reruns once, which is all that is needed.
"""
url = self.search_url(query)
self._ensure_browser()
assert self._browser is not None
print(f"Navigating to {url} ...")
browser_open_url(self._browser, url)
print(f"Waiting {RENDER_TIMEOUT}s for the client-rendered page ...")
time.sleep(RENDER_TIMEOUT)
self._scroll_to_load()
time.sleep(SCROLL_PAUSE)
return dom_element_html(self._browser, "body", outer_html=True)
@staticmethod
def _save_sample(html_path: Path, query: str, html: str) -> None:
if html_path.exists():
html_path.replace(SAMPLE_HTML_BACKUP_PATH)
print(
f"Backed up previous {html_path.name} to "
f"{SAMPLE_HTML_BACKUP_PATH.name}"
)
html_path.write_text(html, encoding="utf-8")
SAMPLE_META_PATH.write_text(json.dumps({"query": query}), encoding="utf-8")
print(f"Saved {len(html)} bytes to {html_path}")
def scrape_search(
self,
query: str,
*,
html_path: Path | None = None,
force: bool = False,
) -> TikTokSearchData:
"""
Return parsed search data for query.
If html_path exists and force is False, parse the local file (no
browser). Otherwise fetch via the browser, save to html_path, then parse.
Raises FirstRunWarmupError when a live fetch returns 0 videos, so the
existing sample.html is left untouched.
"""
if html_path and html_path.exists() and not force:
html = html_path.read_text(encoding="utf-8")
return parse_search_page(html, query=query)
html = self.fetch_search_html(query)
data = parse_search_page(html, query=query)
if data.result_count == 0:
if html_path:
failed_path = html_path.with_suffix(".failed.html")
failed_path.write_text(html, encoding="utf-8")
print(f"Saved the failed fetch to {failed_path} for inspection")
raise FirstRunWarmupError(FIRST_RUN_MESSAGE)
if html_path:
self._save_sample(html_path, query, html)
return data
def close(self) -> None:
if self._browser:
kill_browser(self._browser)
self._browser = None
def __enter__(self) -> "TikTokClient":
return self
def __exit__(
self,
exc_type: type[BaseException] | None,
exc: BaseException | None,
tb: object | None,
) -> None:
self.close()"""
TikTok search result scraper.
Default: parse sample.html (local, no browser). Live fetches run **headless**
by default; pass --headed to open a visible window (needs a display or Xvfb).
Usage:
uv run python main.py # offline, parse sample.html
uv run python main.py --fetch # live fetch (headless) + parse
uv run python main.py --fetch --headed # live fetch in a window
uv run python main.py --query "new trend" --fetch # fetch a different query
The first live fetch on a fresh browser profile returns 0 videos on purpose
(it only warms cookies/caches). Rerun the same command once and it works.
"""
import argparse
import json
import sys
from api import FirstRunWarmupError, TikTokClient
from constants import DEFAULT_QUERY, SAMPLE_HTML_PATH, SAMPLE_META_PATH
def print_separator(char: str = "=", length: int = 60) -> None:
print(char * length)
def _count(value: int | None) -> str:
return str(value) if value is not None else "-"
def _cached_query() -> str:
"""Query that produced the bundled/cached sample.html (default fallback)."""
if SAMPLE_META_PATH.exists():
try:
meta = json.loads(SAMPLE_META_PATH.read_text(encoding="utf-8"))
query = meta.get("query")
if isinstance(query, str) and query:
return query
except (OSError, ValueError):
pass
return DEFAULT_QUERY
def main() -> None:
parser = argparse.ArgumentParser(description="TikTok search scraper")
parser.add_argument(
"--query",
default=DEFAULT_QUERY,
help=f"search query (default: {DEFAULT_QUERY!r}); needs --fetch if it "
"differs from the cached sample",
)
parser.add_argument(
"--fetch",
action="store_true",
help="Force a live browser fetch even if sample.html exists",
)
parser.add_argument(
"--headed",
action="store_true",
help="Open a visible browser window instead of headless (needs a "
"display; on a server use `xvfb-run -a ... --headed`)",
)
args = parser.parse_args()
html_path = SAMPLE_HTML_PATH
cached_query = _cached_query()
if not args.fetch and args.query != cached_query:
print(
f"No downloaded cache for query {args.query!r} "
f"(the cached sample is for {cached_query!r}).\n"
"Add --fetch to download it:\n"
f" uv run python main.py --query {args.query!r} --fetch\n"
"Expect the first run to return 0 videos - if so, rerun the exact "
"same command once. This is normal on a fresh browser profile.",
file=sys.stderr,
)
raise SystemExit(2)
if html_path.exists() and not args.fetch:
print(f"Using cached {html_path} ({html_path.stat().st_size} bytes)")
else:
print("Fetching from browser ...")
try:
with TikTokClient(headless=not args.headed) as client:
data = client.scrape_search(
args.query, html_path=html_path, force=args.fetch
)
except FirstRunWarmupError as exc:
print(f"\n{exc}", file=sys.stderr)
raise SystemExit(1) from exc
except RuntimeError as exc:
print(f"\nBrowser startup failed:\n{exc}", file=sys.stderr)
raise SystemExit(1) from exc
print()
print_separator()
print(f"SEARCH: {data.query!r} ({data.result_count} videos)")
print_separator()
for i, video in enumerate(data.videos, 1):
print(f"\n {i:>2}. @{video.author_username} ({video.author_nickname})")
print(f" url : {video.video_url}")
print(f" video id : {video.video_id}")
print(f" desc : {(video.description or '-')[:90]}")
print(
f" likes : {_count(video.likes)}"
f" comments: {_count(video.comments)}"
f" shares: {_count(video.shares)}"
f" plays: {_count(video.plays)}"
)
print(f" posted : {video.create_time_text or '-'}")
if video.hashtags:
print(f" hashtags : {' '.join(video.hashtags)}")
if video.cover_url:
print(f" cover : {video.cover_url[:70]}...")
print()
print_separator()
print("Done.")
if __name__ == "__main__":
main()browser/ - the browser control layerThese files (browser.py, actions.py, utils.py) open Brave/Chrome with a debugging port, connect over CDP, navigate, scroll, and grab the finished HTML. They ship ready to use.
Your ~/tiktok folder should contain exactly this:
tiktok/
├── index.html
├── TUTORIAL.md
├── pyproject.toml
├── uv.lock
├── constants.py
├── api.py
├── tiktok_parse.py
├── fetch_sample.py
├── main.py
├── sample.html
├── browser/
│ ├── init.py
│ ├── browser.py
│ ├── actions.py
│ └── utils.py
├── models/
│ ├── init.py
│ ├── browser.py
│ └── search.py
└── data/
└── .gitkeep
Some extra files appear only after you run things - they are safe to ignore or delete:
sample.html.bak # previous good sample, kept when a new fetch succeeds
sample.failed.html # the error page from a failed (0-video) fetch
sample.meta.json # which query produced sample.html
browser/browser.log # browser messages, useful when something goes wrong
You're done! If every file is in place and you ran uv sync, the scraper works. Run it anytime with:
cd ~/tiktok
uv run python main.pyQuestions? Double-check file names and paths - most issues are a missed step above or a typo.
Katy Salgado - October 30, 2025
Why Residential IP Intelligence Services Are Highly Inaccurate?
Katy Salgado - November 13, 2025
Why Unmetered Proxies Are Cheaper (Even With a Lower Success Rate)
Katy Salgado - November 27, 2025
TCP OS Fingerprinting: How Websites Detect Automated Requests (and How Proxies Help)
Katy Salgado - December 15, 2025
Analyzing Competitor TCP Fingerprints: Do Their Opt-In Networks Really Match Their Public Claims?