# How ReClip's Frontend Deduplicates URLs: A Deep Dive into the parseUrls Function

> Learn how ReClip's frontend deduplicates URLs using the JavaScript Set data structure in the parseUrls function. Discover the efficient client-side URL management.

- Repository: [Avery Gan/reclip](https://github.com/averygan/reclip)
- Tags: internals
- Published: 2026-09-03

---

**ReClip removes duplicate URLs client-side using a JavaScript `Set` data structure in the `parseUrls` function located in [`templates/index.html`](https://github.com/averygan/reclip/blob/main/templates/index.html).**

ReClip is an open-source video downloading tool built with Flask and vanilla JavaScript. To prevent redundant API calls and reduce server load, the application performs **URL deduplication entirely in the browser** before sending any data to the backend. This article examines the exact mechanism, implementation details, and practical reuse patterns from the source code.

## The Core Deduplication Logic in templates/index.html

The deduplication happens in [`templates/index.html`](https://github.com/averygan/reclip/blob/main/templates/index.html) within the `parseUrls` function. Here is the actual implementation:

```javascript
function parseUrls(text) {
  // Split on spaces, commas, or newlines → trim → keep only http URLs
  // Then deduplicate with a Set
  return [...new Set(text.split(/[\s,]+/).map(u => u.trim())
                .filter(u => u.startsWith('http')))];
}

```

This single-line function performs four sequential operations:

1. **Split** — `text.split(/[\s,]+/)` breaks input on whitespace, commas, or newlines
2. **Trim** — `.map(u => u.trim())` removes leading/trailing whitespace from each token
3. **Filter** — `.filter(u => u.startsWith('http'))` keeps only valid URL strings
4. **Deduplicate** — `new Set(...)` automatically discards duplicate values, then `[...Set]` spreads back to an array

Because a **JavaScript `Set`** only stores unique values by definition, any URL appearing multiple times in user input is reduced to a single entry. This deduplication happens **before any network request**, ensuring [`app.py`](https://github.com/averygan/reclip/blob/main/app.py) receives a clean list of distinct URLs.

## Why Client-Side Deduplication Matters

ReClip's architecture separates concerns deliberately:

- **Frontend** ([`templates/index.html`](https://github.com/averygan/reclip/blob/main/templates/index.html)) — validates, normalizes, and deduplicates URLs
- **Backend** ([`app.py`](https://github.com/averygan/reclip/blob/main/app.py)) — handles `/api/info`, `/api/playlist`, and `/api/download` endpoints with no duplicate-handling overhead

This design choice reduces server load and prevents wasteful processing of repeated video URLs. According to the ReClip source code, the backend assumes all incoming URL lists are already unique.

## Practical Reuse Examples

The `parseUrls` pattern can be extracted and reused throughout your application or in external scripts.

### Example 1: Direct Function Call

```javascript
const rawInput = "https://youtu.be/abc https://youtu.be/abc https://vimeo.com/xyz";
const uniqueUrls = parseUrls(rawInput);
console.log(uniqueUrls);
// Output: ["https://youtu.be/abc", "https://vimeo.com/xyz"]

```

### Example 2: Node.js Module Export

```javascript
// utils/urlDedup.js
export function dedupUrls(text) {
  return [...new Set(text.split(/[\s,]+/).map(u => u.trim())
                .filter(u => u.startsWith('http')))];
}

// usage
import { dedupUrls } from './utils/urlDedup.js';
const urls = dedupUrls("<user-provided string>");

```

### Example 3: Inline Script Block

```html
<script>
  const input = document.getElementById('urls').value;
  const unique = [...new Set(input.split(/[\s,]+/).map(u=>u.trim()).filter(u=>u.startsWith('http')))];
  console.log(unique);
</script>

```

All three variants use the same core technique: **convert the URL array to a `Set` and back to an array**.

## Key Source Files for Reference

| File | Purpose |
|------|---------|
| [`templates/index.html`](https://github.com/averygan/reclip/blob/main/templates/index.html) | Contains the `parseUrls` function that performs frontend URL deduplication |
| [`app.py`](https://github.com/averygan/reclip/blob/main/app.py) | Flask backend; receives already-deduplicated URLs from the frontend |
| [`README.md`](https://github.com/averygan/reclip/blob/main/README.md) | Documents the automatic URL deduplication feature |

## Summary

- ReClip's **URL deduplication** occurs entirely in [`templates/index.html`](https://github.com/averygan/reclip/blob/main/templates/index.html) via the `parseUrls` function
- The implementation leverages **JavaScript's native `Set` data structure** for automatic uniqueness
- Processing happens **client-side**, reducing backend load and eliminating redundant API calls
- The pattern is portable: extract `parseUrls` or inline the `[...new Set(...)]` transformation anywhere URL cleaning is needed

## Frequently Asked Questions

### What happens if I paste the same URL twice in ReClip?

ReClip's `parseUrls` function automatically removes the duplicate. The `Set` constructor in [`templates/index.html`](https://github.com/averygan/reclip/blob/main/templates/index.html) ensures only unique values survive, so the backend receives each URL exactly once regardless of input repetition.

### Does ReClip deduplicate URLs on the server side?

No. The [`app.py`](https://github.com/averygan/reclip/blob/main/app.py) backend assumes deduplication is complete. All filtering happens in the browser through the `parseUrls` function in [`templates/index.html`](https://github.com/averygan/reclip/blob/main/templates/index.html). This design keeps server logic simpler and reduces computational overhead.

### Can I use ReClip's deduplication logic in my own project?

Yes. The `parseUrls` pattern is a standard JavaScript idiom: split input, filter valid URLs, then wrap with `[...new Set(...)]`. Adapt the exact implementation from [`templates/index.html`](https://github.com/averygan/reclip/blob/main/templates/index.html) or the module export example above for your own URL-processing needs.

### Why does ReClip use `startsWith('http')` instead of proper URL validation?

The filter prioritizes speed and simplicity over strict validation. The `startsWith('http')` check catches both `http://` and `https://` URLs while filtering out empty strings and obvious non-URLs. Full validation is deferred to downstream processing or user feedback.