My experience on my daily works... helping others ease each other

Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

Monday, June 1, 2020

Reading entire URL content is really easy using R

In my good old days, reading the entire content of a website is not easy. The process of web scraping and getting the required data requires lots of programming and a few tools. A friend of mine even developed and sold the tool which he called it (during the development) as myrobot. He developed using PHP.

Now, it is much easier and one of the many ways is using R.

Here are the steps (which requires you to write ONLY two lines of code)

  1. Connect to the website using URL command
    con <- url ([the website url], “r”)
  2. Then, read the website
    x <- readLines(con)
  3. Do whatever you wish with the data. In this example, I print out the head of the website and also copy the whole content to a file.
    head(x)
    dput(x, “readFromUrlExample.html”)

There you go.

Result of the head(x) function
Snapshot of the content of the file copied into readFromUrlExamplehtml

The sample source code can be retrieved at 

https://github.com/masteramuk/LearnR-Coursera/blob/master/sample-ReadFromUrl.R

Share:

Friday, May 29, 2020

Solving Committing Issue between R Studio and Github


Solving Committing Issue between R Studio and Github

In the normal development process, you will create a repo (the repo in this article is located at Github), followed by the cloning process or download as full directory into your localhost. It is much easier and straightforward. There won’t be any issues especially if your scrum master or release manager is a well trained person in handling branching, merging, and releasing code using git.


However, in most cases, especially for a full-stack developer who did everything on its own, you may encounter an issue if:

  1. You created a project in your localhost first using R Studio and set Git as your SVN through your project setting
  2. Then you created the repo at the GitHub
  3. Finally, upon ready, you run command to sync with your GitHub

The following is the command that you use/execute and the result of running the command:

% git remote add origin [your GitHub report url]
% git pull origin master
warning: no common commits
remote: Enumerating objects: 3, done.
remote: Counting objects: 100% (3/3), done.
remote: Compressing objects: 100% (2/2), done.
remote: Total 3 (delta 0), reused 0 (delta 0), pack-reused 0
Unpacking objects: 100% (3/3), done.
From [your GitHub report url]
* branch master -> FETCH_HEAD
* [new branch] master -> origin/master
fatal: refusing to merge unrelated histories

and you see the last sentence .. ERROR


Then, based on google, you followed with the following command

% git push -u origin master

and you get the following response (or similar)

To [your GitHub report url]
! [rejected] master -> master (non-fast-forward)
error: failed to push some refs to ‘
[your GitHub report url]'
hint: Updates were rejected because the tip of your current branch is behind
hint: its remote counterpart. Integrate the remote changes (e.g.
hint: ‘git pull …’) before pushing again.
hint: See the ‘Note about fast-forwards’ in ‘git push — help’ for details.

Next, you try to pull again to get the latest branch based on the previous error by running the command to pull again

% git pull origin master

And the result is still not promising 
From [your GitHub report url]
* branch master -> FETCH_HEAD
fatal: refusing to merge unrelated histories

What are you missing or wrongly done? I won’t be able to tell you the missing or wrong steps, but I’m sharing your step to overcoming the problem.


STEPS

  1. Go to you localhost directory where you created the project
  2. In that directory, you should find a file name .gitignore & folder .git
  3. Delete both by running rm -fr (if you are using windows, the command might be different)
  4. Next, init your project file again by running the command git init. You shall see the following message appear after executing the command — Initialized empty Git repository in [your project path]
  5. Followed by adding the remote repo by running the command git remote add origin [your GitHub repo url]
  6. The followed by git add . (make sure there is ‘.’ at the end of the command). It tells the git to add all directories in the remote repo to your local.
  7. Followed by git pull origin master. If succeed, you shall be able to see the following result:
    remote: Enumerating objects: 3, done.
    remote: Counting objects: 100% (3/3), done.
    remote: Compressing objects: 100% (2/2), done.
    remote: Total 3 (delta 0), reused 0 (delta 0), pack-reused 0
    Unpacking objects: 100% (3/3), done.
    From [your GitHub repo url]
    * branch master -> FETCH_HEAD
    * [new branch] master -> origin/master
  8. Finally, run git push -u origin master to verify again. You shall see the following result to indicate it is successfully integrated between your local repo and your Github repo and R Studio shall be able to interact perfectly with GitHub.
    Branch ‘master’ set up to track remote branch ‘master’ from ‘origin’. Everything up-to-date

Once you have done all the steps, go ahead to your R Studio and perform Stage -> Commit -> Commit Message -> Push to complete the process. Refresh your Github page and you shall see all of your local files at your GitHub repo.

If you find this useful, you can buy me a coffee :) @ https://www.buymeacoffee.com/masteramuk

Share:

Saturday, April 18, 2020

Tableau Public - Apple Mobility Data and Dark Mode

I've been working on 2 visualizations. 1 is based on Apply Mobility Data and the other is based on edited Sample Superstore.

Check it out here.

Mobility Data based on Apple's mobility data.



Grid on Dark Mode based on Tableau Sample Superstore (edited version)

All visual is available at https://public.tableau.com/profile/nurul.haszeli.ahmad#!/

All dataset is available at data.world @ https://data.world/haszeliahmad/data-analysis
Share:

Thursday, April 2, 2020

Data Analysis - Use it !!

I was reading many developer's site and chat (telegram and whatsapp) when the government stated that they are looking for an app similar to Singapore apps to track the close contact of the Covid-19 positive patient.

In Singapore, they are using a technology which I presume is Bluetooth to ping close contact within the radiant of the tech and capture necessary data which then used to determine the contact and request them to perform screening. Here are some of the news:

  1. https://www.thestar.com.my/tech/tech-news/2020/03/20/covid-19-singapore-launches-contact-tracing-mobile-app-to-track-coronavirus-infections
  2. https://www.pymnts.com/coronavirus/2020/app-lets-singapore-track-virus-patients-movements/
  3. https://www.nst.com.my/news/nation/2020/03/578445/smartphone-app-track-contacts-covid-19-patients
  4. https://asia.nikkei.com/Spotlight/Coronavirus/Singapore-urges-citizens-to-sign-up-for-COVID-19-tracking-app


And the apps is available on Google Play and Apple Store

  1. https://play.google.com/store/apps/details?id=sg.gov.tech.bluetrace&hl=en
  2. https://apps.apple.com/sg/app/tracetogether/id1498276074

And as this article is written, there are many groups including international are coming with various hackathons for apps that can be used to track all COVID-19 patients and their close contact.

In Malaysia, since the announcement, many had gather groups to develop apps.

From my perspective, why must we reinvent the wheel? Why we need to develop many apps when we already have few that are potentially be used for it. 

For instance, D'scover by Favioriot was developed for a user to explore whatever the user likes but also close contact that uses the apps. I believed they can just tweak the apps to get close contact for COVID-19 and it is much faster than building and testing new apps. (by the way, this is not promoting them and I don't gain anything from it :))

Not just that apps, people have been using Google Maps, Waze, Grab, etc. All these apps collected millions of data and one of these data is people's location and whereabouts. On top of that, all telcos do have their customer's data location and track their movement. I attended the Big Data conference by Bigit a few years ago where one of the telcos presented their data analysis. They have shown the heatmap of their user and based on the communication tower.

I even had a discussion with a few telcos when they approach us (my previous company) to provide their services and wish for data sharing. I do request to have a set of data of their customer whereabouts too to ensure we can provide efficient services at the moment the customer approaches our station or hub, or at the time they are supposed to do so.

These data can be utilized to find close contact with COVID-19 patients. From these data, we know where they go, their ride and whom the came across with or pass by. Of course, these data are secured by all those companies for customer's safety and PDPA compliance. But, for the sake of government and to combat COVID-19, they can request minimal information limited to the phone number to call the respective COVID-19 contact. 

You just need a group of data scientists and data engineers to focus on the massaging and provide the relevant info to the government fast and secure. That's all :)

Don't REINVENT the wheel. Used It and MAXIMIZE the POTENTIAL.

* My personal opinion based on experience. Agreed to disagree :)


Share:

Tuesday, February 25, 2020

Running Pentaho Data Integration in Mac OS Catalina version 10.15 and above


If you haven’t download the application, you may access here https://community.hitachivantara.com/s/article/downloads

Running Pentaho Data Integration @ Spoon in Windows or Linux should be straight-forward. Either click on spoon.bat or spoon.sh or the Data Integration app icon.
Running Data Integration from the install folder
However, for Mac OS, especially with the latest version Catalina which only allowed certified and trusted 64-bit application to run, running Data Integration will be troublesome despite the security for the app was disabled and allowed for the application to run. 





Some of the guidance on enable security and allow the untrusted app to run are as follow:
  1. https://edpflager.com/?p=3571
  2. https://andres.jaimes.net/1388/running-pentaho-spoon-on-mac-osx/
I have tried both and after enabling the apps, click on the Data Integration icon still does not work and the application still did not run. 

Lastly, I had to try to run it through the terminal. For OSX Catalina and above, instead of normal bash, Apple brings in zsh and the behavior is totally different. Plus, if you are a developer and install many Java JDK, running the Pentaho Data Integration will not be as running the normal command. Here is the step to run on the terminal and it works for all OSX.

1. Open up your terminal


2. Navigate to your install path (where you install or unzip the data integration file)


3. Run ./spoon.sh (for bash or old scripting, running spoon.sh shall work without ‘./’)

4. If you run into an error such as JDK or java runtime error like below, do not panic. This maybe your current java JDK is set to be higher than supported by the application.

(base) MyMek @ MyEpal data-integration % ./spoon.sh
OpenJDK 64-Bit Server VM warning: Ignoring option MaxPermSize; support was removed in 8.0
-Djava.endorsed.dirs=/Users/masteramuk/Documents/Apps/data-integration/system/karaf/lib/endorsed is not supported. Endorsed standards and standalone APIs
in modular form will be supported via the concept of upgradeable modules.
Error: Could not create the Java Virtual Machine.
Error: A fatal exception has occurred. Program will exit.

5. First, check your JDK version. Open a new terminal as the command need to be executed from the base (unless you set the JDK in your profile which may cause problem to run multiple JDK later). Run command java -version. You shall have similar to below. In my laptop, the current JDK version is set to JDK version 10.

(base) MyMek @ MyEpal data-integration % java -version
openjdk version "10.0.2" 2018-07-17
OpenJDK Runtime Environment AdoptOpenJDK (build 10.0.2+13)
OpenJDK 64-Bit Server VM AdoptOpenJDK (build 10.0.2+13, mixed mode)

6. Make sure you have multiple java JDK installed if you need to use the existing JDK version for your ‘other’ development. Refer https://www.devdungeon.com/content/install-multiple-jdk-windows-java-development to install multiple JDK. Refer here https://www.jenv.be/ to install jenv command tool.

7. For my laptop, I have a few versions of JDK and using jenv, I can set the selected JDK for global or local (only applied to the folder where we run jenv local command). Run jenv version to check on current JDK and jenv versions to list all JDK install. 

(base) MyMek @ MyEpal data-integration % jenv versions
  system
  1.8
  1.8.0.232
* 10.0 (set by /Users/masteramuk/Documents/Apps/data-integration/.java-version)
  10.0.2
  13.0
  13.0.1
  9
  openjdk64-1.8.0.232
  openjdk64-10.0.2
  openjdk64-13.0.1
  openjdk64-9

As of this manual written, Pentaho Data Integration @ Spoon supports up to JDK 1.8.

8. Use jenv to change the jdk to 1.9. Run command jenv local [JDK number]. In this example, I execute jenv local 1.8.0.232.
9. Finally, run again ./spoon.sh. The application shall run.

Running


Opening the application

The main screen

Share:

Wednesday, January 1, 2020

Tableau For Beginner

I'll be publishing an ebook on Visualizing using Tableau. To those interested, please PM ya. Here is the front page.




Share:

Friday, September 13, 2019

Improving JP with Improvised Prediction Model

Yesterday I wrote on low ridership (read here), of which out of many factors, information availability for journey planning contributes 12% from overall factors. However, the value is based on a survey on one location; that is Penang. Nonetheless, I believed, information availability is the key importance for a smooth journey planning, no matter of the services used or the impact to ridership. The reason for this is based on comments in Google Play for various Journey Planner such as Moovit, Transit Apps, SWIVL, SITS, etc, whereby many user stated their frustration on the accuracy of information displayed by the apps.

Beside information availability being a vital role for riders to plan their journey and services to use, based on the articles referred to in the post, the information must be also reliable, accurate and at real-time (or at least near real-time). If you received an information that was accurate a few minutes ago, there is a probability of the information to be inaccurate at the time of view or receive resulting in the inaccurate plan and action.

How to improve information accuracy, no matter how and when the information arrives at the user?
I did a quick review too on the following articles:

  1. https://www.papercast.com/insights/predict-accurate-bus-arrival-journey-times/
  2. https://core.ac.uk/download/pdf/82293981.pdf
  3. https://escholarship.org/uc/item/51t364vz
  4. http://gamma.cs.unc.edu/TROUTE/
  5. https://www.researchgate.net/publication/274028208_Multimodal_Public_Transit_Trip_Planner_with_Real-Time_Transit_Data
  6. https://repositorio-aberto.up.pt/bitstream/10216/6817/2/26915.pdf
  7. https://jungleworks.com/predicting-accurate-arrival-time/
  8. https://www.researchgate.net/publication/332342499_Survey_of_ETA_prediction_methods_in_public_transport_networks
  9. https://dl.acm.org/citation.cfm?id=3219819.3219874
  10. https://pdfs.semanticscholar.org/8c95/f20cd049e5f0d35466544958631e3e10c258.pdf
  11. https://www.papercast.com/wp-content/uploads/2017/06/Papercast_A4_Better-ETA_2017.pdf
  12. https://datascience.stackexchange.com/questions/10301/how-to-predict-eta-using-regression
  13. https://ieeexplore.ieee.org/abstract/document/1212964
  14. https://www.tandfonline.com/doi/abs/10.1080/15472450600981009
  15. https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1467-8667.2004.00363.x
  16. https://journals.sagepub.com/doi/abs/10.3141/1666-12
  17. https://link.springer.com/chapter/10.1007/978-981-13-3393-4_29
  18. https://patents.google.com/patent/US10254119B2/en
  19. https://arxiv.org/abs/1904.05037
  20. https://patents.google.com/patent/US20190130260A1/en
  21. https://www.tandfonline.com/doi/abs/10.1080/19427867.2017.1366120
  22. https://link.springer.com/article/10.1007/s12652-019-01198-1
  23. https://arxiv.org/abs/1904.03444
  24. https://patents.google.com/patent/US20190051154A1/en
  25. https://ieeexplore.ieee.org/abstract/document/8691701

Based on the articles above, below is the gist of the findings:

  1. Reliable and accurate information at real-time (or near real-time) is critical for smooth journey planning
  2. Recent technology (BDA, AI, ML & IoT) has resulted in many new algorithms to predict accurate ETA to be used in JP
  3. The most recent is KNN which requires lots of historical data and real-time tracking for accurate prediction


Recommendation:

  1. To research and try-n-error all the algorithms to find the best algorithm to predict accurate ETA to be used in the Malaysian environment
  2. To define the best algorithm based on time of request (peak or non-peak)
  3. To come out with a new algorithm and flow to ensure the ETA for JP is 95% accuracy during non-peak and 90% accuracy during peak.


Share:

Saturday, November 3, 2018

Big Data - we forgot to clean our data

Recently I have been working with lots of data coming from various business area such as maintenance, financial transaction, etc., and I found an interesting thought from much top management and young leaders whom don't have enough experience handling data from the source up till visualization.

The first thing comes out from their mind was can it be done in a few hours (or some of them thought it was in a blink of the eyes or in split seconds). Normal question was "When can we see the report or chart? Can we see it tomorrow" and the worst I get "I want it to be ready by today before noon".

Image result for unclean dataMy first reaction was WTF!! (but I won't say it loud). I will normally negotiate with them as most of them don't know the process and the data that they requested. Most of them are easily attracted/amazed by superb visualization presentation by Visualization Tool's Marketing team. The thought everything can be done easily as those marketing people said. They just forgot that in a business presentation, the data set used by those marketing people are prepared and cleaned before being utilized in the tool. For example, Tableau's presentation will always use Sales data for their sample.

It is true that many of visualization tools nowadays are capable to process and display any kind of data. With certain skills, you can massage and clean your data on the fly. I've done that and I know it can be done.  BUT... surely at a cost which from my perspective, it is no beneficial at all.

Why do I say so?


  1. You cannot guarantee that the data you read is 100% clean. You might need to do lots of conversions, data massaging, replacements and calculations. This will definitely incur additional processing power and time during report population. I've come across with many data which either irrelevant, unclean (character in a supposed to be numerical column, date define as string, etc), or contain null values.
  2. You may need to perform lots of table joint or union which can cause your report server or tool to be resource hungry.
  3. You need to understand the data too. Each column and how it shall be visualize must be understood before you can present it correctly.


Related image

That's all from me..Adios









Share:

Sunday, February 18, 2018

List of Data Sciences and Machine Learning usefull link

Credit to: Shivam Panchal
As published by Shivam at Linked @ https://www.linkedin.com/pulse/data-science-machine-learning-beginners-path-shivam-panchal/

Platforms:

  1. What Is Hadoop? Hadoop Tutorial For Beginners https://youtu.be/n3qnsVFNEIU
  2. What is Apache Spark? The big data analytics platform explained http://www.techworld.com.au/article/629920/what-apache-spark-big-data-analytics-platform-explained/
  3. Apache Spark Tutorial: ML with PySpark https://www.datacamp.com/community/tutorials/apache-spark-tutorial-machine-learning
  4. A Beginner's Guide To Apache Pig https://hortonworks.com/tutorial/beginners-guide-to-apache-pig/
  5. Realtime Event Processing in Hadoop with NiFi, Kafka and Storm https://hortonworks.com/tutorial/realtime-event-processing-in-hadoop-with-nifi-kafka-and-storm/

Math:

  1. A Deep Dive Into Linear Algebra https://www.khanacademy.org/math/linear-algebra
  2. An Introduction to Combinatorics & Graph Theory https://www.whitman.edu/mathematics/cgt_online/cgt.pdf

Tools & Framework:

  1. TensorFlow Tutorial – Deep Learning Using TensorFlow https://youtu.be/yX8KuPZCAMo
  2. A 6-part introduction to the MXNet API https://becominghuman.ai/an-introduction-to-the-mxnet-api-part-1-848febdcf8ab
  3. Keras Tutorial: The Ultimate Beginner's Guide to Deep Learning in Python https://elitedatascience.com/keras-tutorial-deep-learning-in-python

Data Visualization:

  1. Building Python Data Apps with Blaze and Bokeh https://youtu.be/1gD9LMqREDs
  2. Matplotlib Tutorial: Python Plotting https://www.datacamp.com/community/tutorials/matplotlib-tutorial-python
  3. Python Bokeh Tutorial - Creating Interactive Web Visualizations https://youtu.be/Mz1AXUE0nR4

Concepts:

  1. Simple Linear Regression https://onlinecourses.science.psu.edu/stat501/node/250
  2. Simple and Multiple Linear Regression in Python https://medium.com/towards-data-science/simple-and-multiple-linear-regression-in-python-c928425168f9
  3. Linear Regression in R https://www.tutorialspoint.com/r/r_linear_regression.htm
  4. An Introduction To Logistic Regression http://ufldl.stanford.edu/tutorial/supervised/LogisticRegression/
  5. Building A Logistic Regression in Python, Step by Step by Susan Li https://medium.com/towards-data-science/building-a-logistic-regression-in-python-step-by-step-becd4d56c9c8
  6. Supervised and Unsupervised Machine Learning Algorithms https://machinelearningmastery.com/supervised-and-unsupervised-machine-learning-algorithms/
  7. 6 Easy Steps to Learn Naive Bayes Algorithm (with codes in Python and R) https://www.analyticsvidhya.com/blog/2017/09/naive-bayes-explained/
  8. A Tutorial on Support Vector Machines for Pattern Recognition http://www.cs.northwestern.edu/~pardo/courses/eecs349/readings/support_vector_machines4.pdf
  9. A Complete Tutorial on Tree Based Modeling from Scratch (in R & Python) https://www.analyticsvidhya.com/blog/2016/04/complete-tutorial-tree-based-modeling-scratch-in-python/

Python:

  1. A Complete Tutorial to Learn Data Science with Python from Scratch https://www.analyticsvidhya.com/blog/2016/01/complete-tutorial-learn-data-science-python-scratch-2/
  2. NumPy Tutorial: Data analysis with Python https://www.dataquest.io/blog/numpy-tutorial-python/
  3. Scipy Tutorial: Vectors and Arrays (Linear Algebra) https://www.datacamp.com/community/tutorials/python-scipy-tutorial
  4. Python Pandas Tutorial https://www.tutorialspoint.com/python_pandas/
  5. Machine Learning with scikit learn Part 1 & 2 https://youtu.be/2kT6QOVSgSghttps://youtu.be/WLYzSas511I

CS:

  1. A Thorough Overview of Computational Logic https://www.cs.utexas.edu/users/boyer/acl.pdf

Game Theory:

  1. Game Theory - A 3 Part Introduction https://youtu.be/x8gOi7D6QeQ

Statistics:

  1. Correlation & causality https://www.khanacademy.org/math/probability/scatterplots-a1/creating-interpreting-scatterplots/v/correlation-and-causality
  2. Analysis of variance (ANOVA) https://www.khanacademy.org/math/statistics-probability/analysis-of-variance-anova-library
  3. Understanding Hypothesis Tests: Significance Levels (Alpha) and P values in Statistics https://shar.es/1PANrc
  4. Characteristics of Good Sample Surveys and Comparative Studies https://onlinecourses.science.psu.edu/stat100/node/3
  5. Descriptive and Inferential Statistics https://www.thoughtco.com/differences-in-descriptive-and-inferential-statistics-3126224
  6. Intro to Probability Theory https://youtu.be/f9XFM8YLccg
  7. Introduction to Conditional Probability & Bayes theorem for data science https://www.analyticsvidhya.com/blog/2017/03/conditional-probability-bayes-theorem/
  8. Central limit theorem https://www.khanacademy.org/math/statistics-probability/sampling-distributions-library/sample-means/v/central-limit-theorem
Share:

About Me

Somewhere, Selangor, Malaysia
An IT by profession, a beginner in photography

Labels

Blog Archive

Blogger templates