# Parsing HTML with Jsoup

**URL:** https://discourse.processing.org/t/parsing-html-with-jsoup/9577
**Category:** Libraries
**Created:** [March 24, 2019, 12:21pm UTC](https://discourse.processing.org/t/parsing-html-with-jsoup/9577 "2019-03-24T12:21:14Z")
**Posts on this page:** 5
**Page:** 1

<div class="post-metadata">

### Author: ![esc746](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.processing.org/esc746/32/4258_2.png) [@esc746](https://discourse.processing.org/u/esc746)
#### Post date: [March 24, 2019, 12:21pm UTC](https://discourse.processing.org/t/parsing-html-with-jsoup/9577/1 "2019-03-24T12:21:14Z")

</div>

I am attempting to use the Jsoup library to parse HTML but the most basic code does not work.

First, the importer generates this:

```auto
import org.jsoup.*;
import org.jsoup.nodes.*;
import org.jsoup.internal.*;
import org.jsoup.parser.*;
import org.jsoup.safety.*;
import org.jsoup.select.*;
import org.jsoup.helper.*;

```

The code is as follows:

```auto
    String url = "https://en.wikipedia.org/wiki/Main_Page";
    Document doc = Jsoup.connect(url).get();
    print(doc.title());
    Elements newsHeadlines = doc.select("#mp-itn b a");
    for (Element headline : newsHeadlines) {
    	print("%s\n\t%s", headline.attr("title"), headline.absUrl("href"));
    }

```

Error: “Unhandled exception type IOException”.

I have tried this with other URLs, including plain HTTP.

Has anyone successfully used this library or have advice about another?

---

<div class="post-metadata">

### Author: ![GoToLoop](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.processing.org/gotoloop/32/86_2.png) [@GoToLoop](https://discourse.processing.org/u/GoToLoop)
#### Post date: [March 24, 2019, 12:26pm UTC](https://discourse.processing.org/t/parsing-html-with-jsoup/9577/2 "2019-03-24T12:26:59Z")

</div>

- [Forum.Processing.org/two/discussions/tagged/Unhandled](http://Forum.Processing.org/two/discussions/tagged/Unhandled)

---

<div class="post-metadata">

### Author: ![esc746](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.processing.org/esc746/32/4258_2.png) [@esc746](https://discourse.processing.org/u/esc746)
#### Post date: [March 24, 2019, 12:35pm UTC](https://discourse.processing.org/t/parsing-html-with-jsoup/9577/3 "2019-03-24T12:35:05Z")

</div>

OK so the following _does_ work. But my peace of mind is shattered, because why would a try block make the IO error simply vanish? Surely, if the error was there, it should now be reported?

```auto
	String url = "https://en.wikipedia.org/wiki/Main_Page";
	Document doc;

	try {
		doc = Jsoup.connect(url).get();
	} catch (IOException e) {
		e.printStackTrace();
		doc = null;
	}

	if (doc != null) {    
	        print(doc.title());
        	Elements newsHeadlines = doc.select("#mp-itn b a");
	        for (Element headline : newsHeadlines) {
        		print("%s\n\t%s", headline.attr("title"), headline.absUrl("href"));
	        }
	}

```

---

<div class="post-metadata">

### Author: ![jeremydouglass](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.processing.org/jeremydouglass/32/20_2.png) [@jeremydouglass](https://discourse.processing.org/u/jeremydouglass)
#### Post date: [March 25, 2019, 9:48pm UTC](https://discourse.processing.org/t/parsing-html-with-jsoup/9577/4 "2019-03-25T21:48:24Z")

</div>

> [@esc746](#):
>
> my peace of mind is shattered, because why would a try block make the IO error simply vanish? Surely, if the error was there, it should now be reported?

You are misunderstanding.

The error you saw is at compile-time, not run-time. In order for your code to use Jsoup, it **must handle** the exception, otherwise your code is invalid. Adding a try-catch makes the code valid, so you get no compile-time error. You are seeing no runtime error because there never was a runtime error – but there could be, and now your code would handle it as required.

> Java has a feature called “checked exceptions”. That means that there are certain kinds of exceptions, namely those that subclass Exception but not RuntimeException, such that if a method may throw them, it _must_ list them in its throws declaration, say: void readData() throws IOException. IOException is one of those. Thus, when you are calling a method that lists IOException in its throws declaration, you must either list it in your own throws declaration or catch it. [java - Why do I get the "Unhandled exception type IOException"? - Stack Overflow](https://stackoverflow.com/a/2305992/7207622)

---

<div class="post-metadata">

### Author: ![esc746](https://yyz2.discourse-cdn.com/flex036/user_avatar/discourse.processing.org/esc746/32/4258_2.png) [@esc746](https://discourse.processing.org/u/esc746)
#### Post date: [March 25, 2019, 10:08pm UTC](https://discourse.processing.org/t/parsing-html-with-jsoup/9577/5 "2019-03-25T22:08:06Z")

</div>

Thanks. I am unfamiliar with Java but your explanation helps.
