When you build a tab snoozer or a "save for later" tool, you quickly realize something: a browser tab is not just a URL. It is an intention, and it is a reading position.
If someone is two-thirds through a 4,000-word documentation page or an RFC and gets pulled into a meeting, closing the tab feels risky. Reopening the tab tomorrow at the very top of the page forces them to hunt for where they left off.
Naturally, the first instinct as a developer is simple: just record window.scrollY when saving the page, and run window.scrollTo() when reopening it.
In practice, that approach turns out to be both unreliable and a permissions nightmare.
Here is why pixel-based scroll restoration breaks in production, and how we solved it using standard W3C text fragments instead.
The Two Traps of Pixel Scroll Offsets
1. Pixel offsets are fragile
A scroll offset is tied to a specific render pass:
- If the user reopens the tab on an external monitor with a different viewport width, text wraps differently and the offset is wrong.
- If the page uses lazy-loaded images or responsive embeds, layout shifts occur after load. By the time the images finish loading, the pixel position points to empty space.
- If the site updates its stylesheet or layout in the days between saving and reopening, a pixel offset becomes completely meaningless.
2. The Manifest V3 permissions trap
To run window.scrollTo() when a tab reopens from a notification click or extension popup, the extension needs to execute a script inside that reopened tab.
Under Manifest V3, activeTab only grants temporary script execution when the user invokes an action directly on that active tab (such as clicking the extension icon or pressing an extension shortcut on that page).
A background notification click opening a tab via chrome.tabs.create() does not qualify as a user gesture on that new tab. To inject a script there, the extension would have to request <all_urls>:
"Read and change all your data on all websites"
Demanding that permission on the install dialog destroys user trust immediately, especially for a tool that promises to be local-first and privacy-focused.
The Solution: Native Text Directives
Instead of storing coordinates, the browser can anchor to the actual content using W3C Scroll-to-Text Fragments:
https://example.com/article#:~:text=the%20sentence%20you%20were%20reading
When Chromium opens that link, the rendering engine finds the specified text, scrolls to it, and highlights it.
This completely changes the permission model:
- Zero code injected at reopening: The browser itself handles the scroll during regular navigation.
- Zero host permissions required: The extension never touches the reopened page.
- Layout-agnostic: It does not matter if images take three seconds to load or if the viewport width changed. The browser anchors directly to the text node.
How We Implemented Capture with activeTab
While reopening requires no permissions, capturing the visible line at save time does require reading the page. Because saving happens via user action (clicking the toolbar icon or pressing Alt+Shift+Y), activeTab applies cleanly.
Here is a simplified version of the node detection logic from our codebase:
function getVisibleReadingAnchor() {
// Ignore saves near the top of the page
if (window.scrollY < 100) return null;
const viewportHeight = window.innerHeight;
// Scan elements starting around 15% from the top
const targetY = viewportHeight * 0.15;
const targetX = window.innerWidth / 2;
let el = document.elementFromPoint(targetX, targetY);
if (!el) return null;
// Skip headers, navigation, footers, and sticky elements
while (el && el !== document.body) {
const tag = el.tagName.toLowerCase();
const style = window.getComputedStyle(el);
if (['nav', 'header', 'footer', 'aside'].includes(tag) ||
style.position === 'fixed' ||
style.position === 'sticky') {
return null;
}
el = el.parentElement;
}
// Extract clean text from the target node
const text = el?.innerText?.trim();
if (!text || text.length < 25) return null;
// Build the text directive query
const cleanSnippet = encodeURIComponent(text.slice(0, 120));
return `#:~:text=${cleanSnippet}`;
}
Edge Cases Discovered in the Wild
Building this across real websites revealed several interesting quirks:
-
Document order is not reading order: On Wikipedia, the table of contents often appears early in the DOM and sits fixed on screen. Early prototypes anchored almost every Wikipedia article to the word "Contents". Skipping
<nav>and sticky containers is essential. - Short labels versus real reading sentences: A page with lots of short metadata (like an issue tracker or release notes) does not have long paragraphs. Setting a strict paragraph length requirement fails on technical docs. We ended up preferring sentences longer than 25 characters, but falling back to shorter phrases if nothing else existed.
-
Graceful degradation: If an article is edited or removed between saving and reopening, the browser silently ignores the
#:~:text=fragment and loads the standard URL from the top. The user is never worse off for having tried.
Takeaway
Before reaching for broad extension permissions or trying to simulate viewport state with JavaScript coordinates, look at what browser standards already provide. Native text fragments let us deliver exact reading position restoration while keeping our extension completely free of install-time permission warnings.
If you are interested in local-first browser utilities or want to see how this feels in practice, we built this into Latr, an open project focusing on tab reminders, tab grouping, and duplicate cleanup without user accounts or tracking servers.
Have you experimented with text fragments or URL directives in your own projects? Would love to hear your thoughts in the comments.
Top comments (0)