ETC5521 Diving Deeper into Data Exploration: Project, part 1

As per Monash’s integrity rules, these solutions are not to be shared beyond this class.

Author

Prof. Di Cook

Published

October 6, 2025

🎯 Goal

The project is designed to challenge you to conduct an exploratory data analysis. Here we are going to explore corporate financials over 2020-2022 as reported from the Bureau van Dijk OSIRIS Database.

The project represents 20% of your final grade for ETC5521. This is a team assignment, and your team member(s) will be allocated by the teaching staff. The last part of the name of your repo should be your team name, eg project-koala.

📌 Guidelines

  1. Check your team allocation by visiting the Google Sheet. The “teams” sheet has your email address along with the name of your team. Get in contact with your other team member(s), as soon as possible, to decide on what country you’d like to work on. The list of possible choices is in the “datasets” sheet. Put your team name against one of the them, only one. First team in, first choice is yours. At most two teams per country.

  2. One team member will need to accept the GitHub Classroom Assignment through the link in moodle using a GitHub Classroom compatible web browser. This will generate a private GitHub repository that can be found at https://github.com/etc5521-2025. You will need to give the other team member(s) access, and change the name to match your team name: The last part of the name of your repo should be your team name, eg if your team name is koala, then the repo name should be project-koala. Your GitHub project repo should contain the file project_part1.html, README.md, report.qmd, project_work_diary.csv, P1.Rproj, and .gitignore.

  3. The diary needs to be updated after every meeting, and after any period of individual work on the project. Working as a team is a skill, needing management of individual strengths, working habits and personality differences, with an end-product that is better than any member can do on their own. The diary should include individual work effort as well as team efforts. Note that marks will be deducted for lack of collaborative effort, by any individual member, and could be as much as 100% deduction.

  4. The first discussions (recommended where you discuss strategy for approaching the project) with your team members needs to be conducted and recorded with zoom and where members are on camera. About 5-10 minutes is sufficient. The recording needs to be uploaded to Moodle.

  5. For the final submission knit the .qmd file and push the resulting .html files to your GitHub repo. You will also need to provide links to any Generative AI conversation you employed in arriving at your solution. Note that marks are allocated for overall grammar and structure of your final report.

  6. We will check your git commit history. Every team member should have substantial and roughly equal in number of contributions to the repo with consistent commits over time.

  7. Refer back to relevant lecture notes and tutorial exercises in developing your solutions. The purpose is to demonstrate what you have learned fron this unit. Only content from this unit can be used. This includes plotting methods, numerical methods, R packages, modeling methods. Every plot, table, and calculation made needs to have a reference to the material, e.g. (Lecture 1, slide 10) or (Reading Week 3). Marks will be deducted for approaches used that have not been discussed in the unit.

  8. You are expected to develop your solutions with your team members, without discussing any details with other class members or other friends or contacts. You can ask for clarifications from the teaching team and we encourage you to attend consultations to get assistance as needed. As a Monash student you are expected to adhere to Monash’s academic integrity policy. and the details on use of Generative AI as detailed on this unit’s Moodle assessment overview.

  9. We expect that there will be 1-2 hours of team meetings, and 7-9 hours of individual work on this project.

  10. If you have a problem with the team work, it needs to be reported here, ideally in a timely manner so the problem can be corrected.

Deadlines:

Due date Turn in
11:45pm Mon Oct 13 Project Repo on GitHub has been created, video discussion uploaded to moodle
11:45pm Mon Oct 20 Final solutions available on repo

🛠️ Exercises

This is data is the OSIRIS from the Bureau van Dijk OSIRIS Database, which contains comprehensive financial and ownership information on public companies, banks, and insurance companies globally. It provides standardised and “as reported” financials, earnings estimates, ownership data, and news. Visit the link in the References section to learn more about the variables reported in this data.

The time period of 2020-2022 covers the pandemic period. So the primary question motivating the analysis is “How did the pandemic affect corporate financials?”

However, remember that exploratory data analysis is about finding interesting and unexpected patterns in data. So while you need to answer the above question in some way, you are also expected to report on several patterns or relationships in this data that are surprising or unexpected.

To help motivate your exploration, have a read through the paper “Boom and Bust of Technology Companies at the Turn of the 21st Century” is in the file hofmann_wickham_cook.pdf.

References

Bureau van Dijk Electronic Publishing. (n.d.). Osiris. Bureau van Dijk Electronic Publishing. https://ezproxy.lib.monash.edu.au/login?url=http://osiris.bvdep.com/ip

Marks

Part Points
Introduction & Data Description 2
IDA 4
EDA 7
Conclusions 3
Collaboration 4
Citations & Generative AI Analysis -5
Reproducibility, Formatting, Spelling & Grammar -5

Note that the negative marks for “Generative AI Analysis”, “Formatting, Spelling & Grammar” correspond to reductions in scores. You can lose up to 3 marks for poor use of the GAI. For example, no use, basic questions only, no link to the script, and no acknowledgment but clearly used. You can lose up to 5 marks if your report is not reproducible, for poorly formatted and written answers. Three marks will be reserved for appropriate GitHub work, accepting the assignment in a timely fashion and consistent and substantive commits.

Rubric

To help you complete in your report, below is a rubric to guide you to what we are expecting:

content text Excellent (HD) Very good (D) Good (C) Satisfactory (P) Unsatisfactory (F)
Introduction Explanation of the problem Motivates and explains the proble to communicate the background. Outline encourages reading of the other sections. Data sources explained, including limitations that might affect possible analysis and conclusions. Explanation of problem of interest is VERY clear and provides information about the background. Details of data sources are provided, with limitations affecting analysis. Explanation of problem is clear and provides information about the background. Explanation of problem is rudimentary and lacks detail. Explanation of problem is unclear and/or not shown. There are no suitable expectations to motivate exploration.
Data description Description of data, number of observations, variables and types. Only include variables used in the analysis. Description of the variables, as organised into tidy form, nicely formatted as a table. Detailed and concise explanation of data being analysed with a comprehensive overview of the methods that might be needed to check and clean it, including reasons. Original data source, and OpenAQ site and software cited. Description of the variables, as organised into tidy form, nicely formatted as a table. Original data source, and OpenAQ site and software site cited. Data description is soundly presented and demonstrates ability to translate tidy data summary into clear data desctipions Data description is reasonably presented and lists basic understanding of translating tidy data form into data descriptions. Description of data is unclear and/or not shown and demonstrates no or little understanding.
IDA Data pre-processing and cleaning Comprehensive data inspection conducted and explained, with perfect choice of plots and summaries. Not too many, not too few. It is clear the clean data is in a state to be further explored. Data inspection conducted and explained well, with good choice of plots and summaries. It is clear the clean data is in a state to be further explored. Data inspection adequate and explained, with some plots and summaries. Data inspection reasonable with some explanation. Data inspection inadequate and insufficient explanation.
EDA Plots and summaries to explore the data, not ncessarily only covering the expectations. Plots and summaries answer questions raised by the initial expectations, are designed well, and give clear answers along with clear reasoning. All are interesting and surprising. Plots include interactive elements to allow mouseover for more information. Tables are interactive to allow re-sorting. Plots and summaries answer questions raised by the initial expectations, are designed well, and provide good answers and reasoning. Most are interesting and surprising. Some plots and tables include appropriate interactive elements to allow mouseover for more information. Plots and summaries mostly answer questions raised by the initial expectations, are provide good answers and reasoning, and some are surprising. Plots and summaries partially answer questions raised by the initial expectations, providing suitable answers. Plots and summaries are not well matched to the initial expectations, with little explanation provided.
Conclusions Concise summary of findings Concise and clear summary organised to be very fun to read. Strongly supprted by clear rationale and reasoning. Concise and clear summary organised to be very readble. Rationale and supporting arguments provided. Concise and clear summary organised to be mostly readable with some rationale and supporting arguments provided. Clear summary organised to be mostly readable. Summary doesn't match expectations or analysis.
Collaboration - Documentation indicates met, talked effectively and regularly, and all individuals clearly contributed based on the diary. Consistent and regular commits by all team members with informative commit messages. Documentation indicates met, talked effectively and regularly, and all individuals clearly contributed based on the diary. Consistent and regular commits by all team members with informative commit messages. Talked effectively, and all individuals clearly contributed based on the diary. Regular commits evenly distributed among team members with informative commit messages. Group met and indiviuals worked on project based on the diary. Evidence of commits by all team members. No diary entries or evidence of team work.
GAI - GAI used effectively, deeply, and explained, script linked to report GAI used effectively, script linked to report GAI used and script linked to report Shallow use of GAI, script linked to report Clearly used but no script linked to report
Reproducibility - No changes needed for report to reproduce exactly as provided. Single change needed for report to reproduce exactly as provided. Just a few changes needed for report to reproduce exactly as provided. Multiple changes needed for report to reproduce exactly as provided. Cannot easily make changes for report to reproduce at all.
Spelling/Grammar Spelling and grammar correct Writing style is exceptional, scholarly and succinct that is free from spelling, grammar and punctuation errors. Writing style is scholarly, free from spelling, grammar and punctuation errors. Writing style is scholarly, but wordy and inconcise. Free from spelling, grammar and punctuation errors. Writing is scholarly and wordy. Contains some grammatical, punctuation and spelling errors. Writing is unscholarly. Many grammatical, punctuation and spelling errors.
References Appropriate citations of sources, literature and software used The appropriate referencing style has been used consistently, with no errors. Includes citations for software used, and data sources. The appropriate referencing style has been used consistently, with very few errors, and includes software used, and data sources. The appropriate referencing style has been used consistently, and only a few citations missing. The appropriate referencing style has been used much of the time, missing some major sources that were clearly used. Material used from external sources without citation.