Aggregating data with Portia

Itsy, Bitsy Spider

Article from Issue 169/2014

Author(s): Tim Schürmann

Are you interested in retrieving stock quotes in machine-readable form off the Internet? No problem: After a few mouse clicks, Portia weaves a command line and wraps the data in JSON format.

The Internet is a treasure trove of useful information, often residing on colorful HTML pages that are not easily extracted and processed. If you want to automate processing of current stock quotes or aggregate news, for example, you need to dismantle the HTML code of news portals such as CNN or Slashdot. This can be pretty ugly work.

Portia, a tool written in Python [1], promises a remedy; its name also refers to a genus of spiders, which would seem to make sense on the World Wide Web. The tool consists of a web application that, with a simple click, allows a user to select stock quotes, messages, and any other desired content. Portia then extracts this data and outputs it in JSON format.

Supported by a supplied web crawler, Portia can also ransack complete websites. As an example, if you need the headings from all Wikipedia articles, you show Portia exactly once where the headline resides on a Wikipedia page. The crawler then traverses the entire website and returns all matching headings in JSON format (see the "Warning" box for more information).

[...]

Use Express-Checkout link below to read the full article (PDF).

Buy this article as PDF

Express-Checkout as PDF

Price $2.95
(incl. VAT)

Buy Linux Magazine

SINGLE ISSUES

Print Issues

Digital Issues

SUBSCRIPTIONS

Print Subs

Digisubs

TABLET & SMARTPHONE APPS

US / Canada

UK / Australia

Support Our Work

Linux Magazine content is made possible with support from readers like you. Please consider contributing when you’ve found an article to be beneficial.

News

Fedora 42 Available with Two New Spins

Fedora , Gnome , Plasma

The latest release from the Fedora Project includes the usual updates, a new kernel, an official KDE Plasma spin, and a new System76 spin.
So Long, ArcoLinux

Linux , open source , Operating Systems

The ArcoLinux distribution is the latest Linux distribution to shut down.
What Open Source Pros Look for in a Job Role

FOSS , open source

Learn what professionals in technical and non-technical roles say is most important when seeking a new position.
Asahi Linux Runs into Issues with M4 Support

Linux , open source

Due to Apple Silicon changes, the Asahi Linux project is at odds with adding support for the M4 chips.
Plasma 6.3.4 Now Available

KDE , Linux , Plasma

Although not a major release, Plasma 6.3.4 does fix some bugs and offer a subtle change for the Plasma sidebar.
Linux Kernel 6.15 First Release Candidate Now Available

Kernel , Linux

Linux Torvalds has announced that the release candidate for the final release of the Linux 6.15 series is now available.
Akamai Will Host kernel.org

Kernel , Linux , Security

The organization dedicated to cloud-based solutions has agreed to host kernel.org to deliver long-term stability for the development team.
Linux Kernel 6.14 Released

Kernel , Linux , Rust

The latest Linux kernel has arrived with extra Rust support and more.
EndeavorOS Mercury Neo Available

EndeavorOS , Linux , Plasma

A new release from the EndeavorOS team ships with Plasma 6.3 and other goodies.
Fedora 42 Beta Has Arrived

Fedora , Linux , Plasma

The Fedora Project has announced the availability of the first beta release for version 42 of the open-source distribution.

Aggregating data with Portia

Itsy, Bitsy Spider

Buy this article as PDF

Buy Linux Magazine

Related content

Subscribe to our Linux Newsletters
Find Linux and Open Source Jobs
Subscribe to our ADMIN Newsletters

Support Our Work

News

Fedora 42 Available with Two New Spins

So Long, ArcoLinux

What Open Source Pros Look for in a Job Role

Asahi Linux Runs into Issues with M4 Support

Plasma 6.3.4 Now Available

Linux Kernel 6.15 First Release Candidate Now Available

Akamai Will Host kernel.org

Linux Kernel 6.14 Released

EndeavorOS Mercury Neo Available

Fedora 42 Beta Has Arrived

Aggregating data with Portia

Itsy, Bitsy Spider

Buy this article as PDF

Buy Linux Magazine

Related content

Subscribe to our Linux Newsletters Find Linux and Open Source Jobs Subscribe to our ADMIN Newsletters

Support Our Work

News

Tag Cloud

Subscribe to our Linux Newsletters
Find Linux and Open Source Jobs
Subscribe to our ADMIN Newsletters