Demystifying Web scraping (BeautifulSoup tutorial)
How to scrape websites using beautiful soup and requests

As a technical writer in the tech industry, I specialize in creating clear, concise, and informative documentation for complex technologies. With a keen eye for detail and a deep understanding of the subject matter, I can distill technical jargon into easy-to-understand language for a wide range of audiences. From user manuals to developer guides, my writing ensures that end-users and stakeholders can easily access and implement the technology they need to succeed. With a strong commitment to quality and a passion for the latest trends and innovations, I am dedicated to providing clients with top-notch technical writing services that help them achieve their goals.
Downloading and installing packages and modules.
In this tutorial, I assume you have a basic knowledge of Python programming.
Before we proceed, I recommend that you peruse my previously written work on the introduction to web scraping
To start with, we need to install the necessary packages and modules. we will install requests, then bs4 and selenium. Head on to your command prompt and install the request using the following PIP (preferred installer program) commands;
pip install requests
After running this, wait a moment for it to automatically download and install the package. Do the same for bs4, and selenium.
pip install bs4
pip install selenium
After installing the above packages, search "download chromedriver" on Google or use this link chromedriver_link. Chrome driver will be required when we start using selenium.
When you're done with downloading chromedriver, head back to your VScode or whichever editor you prefer to use.
The functionality of each installed module.
It'd be necessary to educate you on the difference between the selenium method and the bs4 / request method.
Requests
is a module in Python that helps you send HTTP requests containing a given URL? It does this, without using any browser. Requests can be used to send GET, POST, Update (PATCH and PUT) or DELETE requests. In this tutorial, we're only concerned with the GET method in requests. GET method in requests fetches every detail present in the URL we're scraping, that is, the URL fed into the "get" method. Instead of relaying it to a browser to decipher the HTML code, "requests" brings the raw response directly to our computer.
bs4:
this is a powerful package with the full meaning, "beautiful soup 4". Bs4 is used to extract contents of HTML or XML and organize these contents, usually in a tree mode for easy access. bs4 is considerably powerful when it comes to data parsing from HTML and XML samples. We shall see more from it later.
Selenium.
Selenium is a very powerful tool built mainly for pen testing. It has broad functionality, hence a wide use case. In this tutorial, we're only concerned about the functionality relevant to web scraping.
Using Requests and Bs4 module.
To start with, copy the code below into your Python script or use your Python interpreter or shell to run it.
Getting Response
Note, I chose one of my live projects, a mini website as a case study to use to avoid any legal issues or accusations relating to our actions.
You can see that http://adebola.pythonanywhere.com is the site we are about to scrape and we passed its URL into the get method of the requests module.
We got a response containing the whole HTML details of the page. The get method has many methods that it responds to also, for example, .reason will output "OK" if it was a successful request, else it returns the error message related to the response code gotten. .text outputs the HTML code of the page as string type or Unicode characters. Printing the raw response itself outputs the response code e.g 200, 201, 203 for responses that are ok and successful while other numbers such as 304, 308, 404, 405, 425, 504, and 505 have special error messages associated with them. You can get a good understanding of requests responses from the Mozilla developer docs.
dot Content(.content) and dot text (.text)
".content" and ".text" performs almost the same task, except that .content outputs the byte representation of the page while ".text" returns a string. You can replace raw_response.content with raw_response.text while scraping another site and notice that there's a slight difference in the two outcomes because one returns the bytes representation of the page while the other returns the page content as Unicode characters. Using the type function in Python on both .content and .text will help you understand this concept without hiccups.
Parsing response
We have to parse the response gotten, that is where beautifulSoup commences its operation. Copy and run this code in your editor:
soup in this case represents the parsed version of the raw_request and thus, can be sorted and used to fetch the desired information. A typical example would be to get all the links on the website and label tags on the page.
You should print each of the variables in this code to visualize what each one contains.
you should give the documentation of beautiful soup a good read and explore its possibilities.
You can also check out the list of methods present in beautifulSoup after importing it as bs, by running:
print(dir(bs))
Downloading images.
To get images and save them on your local device, you have to pick out the link to that image, using the same approach we used for links, then send a request and write the image file as bytes.
Continue with the previously written code.
I hope you enjoyed the moment spent here and I also assume that you acquired some new knowledge as you read through. Do realize that we are yet to try our hands on selenium which is a more powerful tool for web scraping.
To be continued......


