When you have a spreadsheet of image links, downloading them by hand is impractical. This script reads a CSV whose first column holds the URL and whose optional second column gives a preferred file name, then fetches every image in parallel using a thread pool. Concurrency is ideal here because image downloads are I/O-bound, so eight workers (configurable) dramatically cut total time versus a sequential loop. Each request streams in 8 KB chunks to keep memory low even for large files, and a per-request timeout prevents one slow or dead URL from stalling the whole batch. When no name is supplied, one is derived from the URL path and sanitized to remove characters that are illegal in file names across operating systems. Failures are caught individually and reported, so one broken link never aborts the run; a final summary tallies successes and failures. Install the only dependency with pip install requests.
Bulk Download Images from a CSV of URLs
Download many images concurrently from a CSV list of URLs, with safe file naming and per-request timeouts.
13 views
Share
Script
Download .pypython
#!/usr/bin/env python3
"""Download images in parallel from a CSV list of URLs.
CSV: first column is the URL; optional second column is a file name.
Dependency: pip install requests
"""
import argparse
import csv
import os
import re
from concurrent.futures import ThreadPoolExecutor, as_completed
from urllib.parse import urlparse
import requests
def read_url_rows(path):
"""Yield (url, name_or_None) tuples from the CSV."""
with open(path, newline="", encoding="utf-8") as f:
for row in csv.reader(f):
if not row or not row[0].strip():
continue
url = row[0].strip()
if url.lower() == "url": # skip a header line
continue
name = row[1].strip() if len(row) > 1 and row[1].strip() else None
yield url, name
def filename_from_url(url, index):
"""Derive a safe file name from the URL path."""
path = urlparse(url).path
base = os.path.basename(path) or f"image_{index}"
base = re.sub(r"[^A-Za-z0-9._-]+", "_", base)
return base
def download_one(url, name, outdir, index, timeout):
"""Download a single image; return (url, status_string)."""
try:
resp = requests.get(url, timeout=timeout, stream=True)
resp.raise_for_status()
fname = name or filename_from_url(url, index)
out_path = os.path.join(outdir, fname)
with open(out_path, "wb") as f:
for chunk in resp.iter_content(chunk_size=8192):
f.write(chunk)
return url, f"OK -> {out_path}"
except requests.RequestException as exc:
return url, f"FAILED ({exc})"
def main():
parser = argparse.ArgumentParser(description="Bulk image downloader.")
parser.add_argument("csv_file", help="CSV with image URLs.")
parser.add_argument("-o", "--outdir", default="images",
help="Output directory (default: images).")
parser.add_argument("-j", "--workers", type=int, default=8,
help="Parallel download workers (default: 8).")
parser.add_argument("-t", "--timeout", type=float, default=20.0,
help="Per-request timeout seconds (default: 20).")
args = parser.parse_args()
os.makedirs(args.outdir, exist_ok=True)
rows = list(read_url_rows(args.csv_file))
ok = 0
fail = 0
with ThreadPoolExecutor(max_workers=args.workers) as pool:
futures = {
pool.submit(download_one, url, name, args.outdir, i, args.timeout): url
for i, (url, name) in enumerate(rows, start=1)
}
for fut in as_completed(futures):
url, status = fut.result()
print(status)
if status.startswith("OK"):
ok += 1
else:
fail += 1
print(f"\nDone. {ok} succeeded, {fail} failed, {len(rows)} total.")
if __name__ == "__main__":
main()How to run
Review the script first, then download or copy it and run it in your environment.
You might also like
bashlinuxIntermediate
Lock inactive users
Finds human accounts with no login activity beyond a chosen threshold and locks them, with a safe dry-run mode by default.
bashlinuxIntermediate
Backup All MySQL Databases with mysqldump
Dumps each MySQL database to its own compressed SQL file.
bashlinuxIntermediate
Bulk Rename Files by Pattern
Renames many files at once using a sed substitution, with dry-run preview.