# How can I web scrape to get images from different websites using processing?

**URL:** <https://discourse.processing.org/t/how-can-i-web-scrape-to-get-images-from-different-websites-using-processing/22052>\
**Category:** Coding Questions\
**Created:** [June 22, 2020, 7:21am UTC](https://discourse.processing.org/t/how-can-i-web-scrape-to-get-images-from-different-websites-using-processing/22052 "2020-06-22T07:21:27Z")\
**Posts on this page:** 3\
**Page:** 1

<div class="post-metadata">

**Author:** ![usr-pyth](https://avatars.discourse-cdn.com/v4/letter/u/c5a1d2/32.png) [@usr-pyth](https://discourse.processing.org/u/usr-pyth)\
**Post date:** [June 22, 2020, 7:21am UTC](https://discourse.processing.org/t/how-can-i-web-scrape-to-get-images-from-different-websites-using-processing/22052/1 "2020-06-22T07:21:27Z")

</div>

Hi! I’m trying to get images from reddit and load them into my processing sketch. How do I do this?

---

<div class="post-metadata">

**Author:** ![kfrajer](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.processing.org/kfrajer/32/196_2.png) [@kfrajer](https://discourse.processing.org/u/kfrajer)\
**Post date:** [June 23, 2020, 12:28pm UTC](https://discourse.processing.org/t/how-can-i-web-scrape-to-get-images-from-different-websites-using-processing/22052/2 "2020-06-23T12:28:14Z")

</div>

First, you need to write a small script to load an image. Here is a starter:

```java
//===========================================================================
// GLOBAL VARIABLES:
PImage solar;

//===========================================================================
// PROCESSING DEFAULT FUNCTIONS:

void settings(){
  size(320,121);
}

void setup(){

  textAlign(CENTER,CENTER);
  rectMode(CENTER);
  
  fill(255);
  strokeWeight(2);
  noLoop();
  
  //SOURCE: https://en.wikipedia.org/wiki/File:Solar-System.pdf
  solar=loadImage("https://upload.wikimedia.org/wikipedia/commons/thumb/6/64/Solar-System.pdf/page1-320px-Solar-System.pdf.jpg");
}

void draw(){
  background(0);
  image(solar,0,0,width,height);
  
}

```

This script is helpful for debugging issues upfront when loading specific images. Can you load images directly (aka manually) from your source site?

After that, you can use [loadStrings](https://processing.org/reference/loadStrings_.html) to load your target size, parse the image tags in the html content and extract the images.

Kf

---

<div class="post-metadata">

**Author:** ![skypickle](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.processing.org/skypickle/32/11863_2.png) [@skypickle](https://discourse.processing.org/u/skypickle)\
**Post date:** [January 22, 2021, 12:19am UTC](https://discourse.processing.org/t/how-can-i-web-scrape-to-get-images-from-different-websites-using-processing/22052/3 "2021-01-22T00:19:52Z")

</div>

This doesnt always work.

For example,  
‘[https://digimon.shadowsmith.com/img/koromon.jpg](https://digimon.shadowsmith.com/img/koromon.jpg)’  
or  
‘[https://www.google.de//images/branding/googlelogo/2x/googlelogo\_color\_272x92dp.png](https://www.google.de//images/branding/googlelogo/2x/googlelogo_color_272x92dp.png)’

dont give any images when substituted in the script but these images clearly exist.
