{"id":14878,"date":"2025-05-27T07:10:51","date_gmt":"2025-05-27T01:10:51","guid":{"rendered":"https:\/\/dtasiagroup.com\/?p=14878"},"modified":"2025-05-27T07:11:48","modified_gmt":"2025-05-27T01:11:48","slug":"toad-data-point-how-do-i-know-that-my-data-is-accurate","status":"publish","type":"post","link":"https:\/\/dtasiagroup.com\/vi\/toad-data-point-how-do-i-know-that-my-data-is-accurate\/","title":{"rendered":"Toad Data Point: How do I know that my data is accurate?"},"content":{"rendered":"<p data-start=\"290\" data-end=\"779\"><img loading=\"lazy\" decoding=\"async\" class=\"aligncenter wp-image-14879 size-full\" src=\"https:\/\/dtasiagroup.com\/wp-content\/uploads\/2025\/05\/4237.Toad-Edge.jpg-800x400x2.jpg\" alt=\"\" width=\"800\" height=\"400\" \/><\/p>\n<p data-start=\"290\" data-end=\"779\">In the world of relational databases, data is typically structured according to normalization principles. The goal is to reduce redundancy and enforce data integrity. This has long been the foundation of database design, especially when data is organized into the Third Normal Form (3NF) or even the Boyce-Codd Normal Form (BCNF). However, as data becomes more normalized, it also becomes less aligned with the format needed for reports, dashboards, and business intelligence applications.<\/p>\n<p data-start=\"781\" data-end=\"1155\">Once you&#8217;ve connected to your data sources, located the relevant information, and created your dataset, the next crucial step is ensuring the accuracy of that data. It\u2019s surprisingly common to discover that your dataset contains inaccuracies, duplicates, or what\u2019s often referred to as \u201cdirty\u201d data. This is where Toad Data Point becomes an essential tool for data analysts.<\/p>\n<p data-start=\"1190\" data-end=\"1646\">In today\u2019s application landscape, much of the data entry process is placed in the hands of end-users. Whether signing up for a website or completing an online form, users are frequently asked to provide personal details like email addresses, phone numbers, and postal codes. However, many users simply input random or incorrect information just to proceed. As a result, databases can quickly become populated with invalid, inconsistent, or misleading data.<\/p>\n<p data-start=\"1648\" data-end=\"1774\">As a data analyst, it\u2019s your job to spot and clean up these issues. But how can you identify bad data quickly and efficiently?<\/p>\n<p>When you\u2019ve created a dataset in Toad Data Point, you can send that dataset to the\u00a0<strong>Data Profiling<\/strong>\u00a0module. The Data Profiling module helps you identify anomalies in your dataset. There are several components to it, as we will investigate below:<\/p>\n<p>To send your dataset for profiling, right click on the dataset results then go to\u00a0<strong>Send To<\/strong>\u00a0|\u00a0<strong>Data Profiling<\/strong>:<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.quest.com\/community\/resized-image\/__size\/320x240\/__key\/communityserver-blogs-components-weblogfiles\/00-00-00-00-37\/5282.Picture1.png\" alt=\" \" \/><\/p>\n<p>When Data Profiling launches, by default it will profile 1,000 rows. If you want to profile your entire dataset you can do so by clicking\u00a0<strong>Edit Profile,<\/strong>\u00a0then selecting\u00a0<strong>All Rows<\/strong>:<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.quest.com\/community\/resized-image\/__size\/320x240\/__key\/communityserver-blogs-components-weblogfiles\/00-00-00-00-37\/1526.Picture2.png\" alt=\" \" \/><\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.quest.com\/community\/resized-image\/__size\/320x240\/__key\/communityserver-blogs-components-weblogfiles\/00-00-00-00-37\/8508.Picture3.png\" alt=\" \" \/><\/p>\n<p>On the\u00a0<strong>Summary<\/strong>\u00a0tab, you will see a breakdown of each column visualized in a chart. The chart, as seen below, is very useful at identifying anomalies. In the example below, I was of the impression that\u00a0<strong>OrderID<\/strong>\u00a0is\u00a0unique, but I can see there are non-unique and repeated rows. To see those repeated rows, I can simply double-click on the orange section of the\u00a0<strong>OrderID<\/strong>\u00a0bar, and the repeated rows will be listed for me.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.quest.com\/community\/resized-image\/__size\/320x240\/__key\/communityserver-blogs-components-weblogfiles\/00-00-00-00-37\/2477.Picture4.png\" alt=\" \" \/><\/p>\n<p>This shows up something interesting \u2013 the Order ID is correctly not always unique. To get the unique value of each row (Primary Key) I need to use\u00a0<strong>OrderID + LineID<\/strong>.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.quest.com\/community\/resized-image\/__size\/320x240\/__key\/communityserver-blogs-components-weblogfiles\/00-00-00-00-37\/4428.Picture5.png\" alt=\" \" \/><\/p>\n<p>Moving to the\u00a0<strong>Statistics<\/strong>\u00a0tab \u2013 if you need statistical information about your data, then this tab is worth the license cost alone. Simply by selecting the column on the left panel, you will instantly get statistical analysis such as median, max, min, average, mode, quartiles, sum, standard deviation, etc. You can also quickly graph the\u00a0<em>Value Distribution<\/em>\u00a0v\u00e0\u00a0<em>Percentiles<\/em>\u00a0for each column. In my opinion, the level of detail this tab gives you is really impressive. The time saved compared to calculating these values manually is significant.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.quest.com\/community\/resized-image\/__size\/320x240\/__key\/communityserver-blogs-components-weblogfiles\/00-00-00-00-37\/6102.Picture6.png\" alt=\" \" \/><\/p>\n<p>&nbsp;<\/p>\n<p>The next tab I want to highlight is\u00a0<strong>Patterns<\/strong>. Very simply, this looks at the patterns of the text\/string columns in your data set. I see customers using this a lot for structured columns such as email address, post code, telephone number.<\/p>\n<p>It will display the Word pattern (are letters, number, punctuation, spaces in use) and the Letter pattern (the order in which letters, numbers, punctuation, spaces are used).<\/p>\n<p>In the example below, it clearly identified anomalies in my email address field, where I have email addresses that contain spaces, making them invalid.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.quest.com\/community\/resized-image\/__size\/320x240\/__key\/communityserver-blogs-components-weblogfiles\/00-00-00-00-37\/8547.Picture7.png\" alt=\" \" \/><\/p>\n<p>&nbsp;<\/p>\n<p>The\u00a0<strong>Language<\/strong>\u00a0tab offers similar information to the\u00a0<strong>Pattern<\/strong>\u00a0tab, this time detailing the character distribution used in text\/string columns. Again, the email address is showing whitespaces in use.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.quest.com\/community\/resized-image\/__size\/320x240\/__key\/communityserver-blogs-components-weblogfiles\/00-00-00-00-37\/2783.Picture8.png\" alt=\" \" \/><\/p>\n<p>&nbsp;<\/p>\n<p>Finally, the Duplicates tab allows you to search for duplicates in the dataset. You simply select the columns you want to search, and it will search for duplicates across all the selected columns. In the below example I\u2019m searching for\u00a0<em>Firstname<\/em>\u00a0v\u00e0\u00a0<em>Surname<\/em>\u00a0duplicates. Since they are \u201cstring\u201d columns, I\u2019m selecting a\u00a0<strong>fuzzy<\/strong>\u00a0search. This tells the Toad to look for similar names, encapsulating potential spelling mistakes on data entry. The below screenshot highlights \u201cJoy Jones\u201d and \u201cJoe Jones\u201d as a potential\u00a0<strong>fuzzy<\/strong>\u00a0duplicate since there is only 1 character difference in their full name. However, I know this isn\u2019t a duplicate, so I\u2019ll add additional columns to make my search more refined.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.quest.com\/community\/resized-image\/__size\/320x240\/__key\/communityserver-blogs-components-weblogfiles\/00-00-00-00-37\/5226.Picture9.png\" alt=\" \" \/><\/p>\n<p>Below I\u2019ve added \u201cDate of Birth\u201d and \u201cEmail Address\u201d, which I\u2019m confident will give me a unique person.<\/p>\n<p><img decoding=\"async\" src=\"https:\/\/www.quest.com\/community\/resized-image\/__size\/320x240\/__key\/communityserver-blogs-components-weblogfiles\/00-00-00-00-37\/1832.Picture10.png\" alt=\" \" \/><\/p>\n<p>Having profiled my dataset, identified anomalies that I need to fix in the source application and\/or revise my query, I can now proceed with confidence knowing that the data I\u2019ve retrieved is accurate.<\/p>\n<p>&nbsp;<\/p>\n<p><strong>About DT Asia<\/strong><\/p>\n<p>DT Asia began in 2007 with a clear mission to build the market entry for various pioneering IT security solutions from the US, Europe and Israel.<\/p>\n<p>Today, DT Asia is a regional, value-added distributor of cybersecurity solutions providing cutting-edge technologies to key government organisations and top private sector clients including global banks and Fortune 500 companies. We have offices and partners around the Asia Pacific to better understand the markets and deliver localised solutions.<\/p>\n<p><strong>\u00a0<\/strong><\/p>\n<p><strong>How we help<\/strong><\/p>\n<p>If you need to know more about Toad Data Point: How do I know that my data is accurate, you\u2019re in the right place, we\u2019re here to help! DTA is Quest Software\u2019s distributor, especially in Singapore and Asia, our technicians have deep experience on the product and relevant technologies you can always trust, we provide this product\u2019s turnkey solutions, including consultation, deployment, and maintenance service.<\/p>\n<p>Click here and here and here to know more:\u00a0<a href=\"https:\/\/dtasiagroup.com\/vi\/quest\/\">https:\/\/dtasiagroup.com\/quest\/<\/a><\/p>","protected":false},"excerpt":{"rendered":"<p>In the world of relational databases, data is typically structured according to normalization principles. The goal is to reduce redundancy and enforce data integrity. This has long been the foundation of database design, especially when data is organized into the Third Normal Form (3NF) or even the Boyce-Codd Normal Form (BCNF). However, as data becomes more normalized, it also becomes less aligned with the format needed for reports, dashboards, and business intelligence applications.<\/p>","protected":false},"author":11,"featured_media":14879,"comment_status":"closed","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[56],"tags":[],"class_list":["post-14878","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-articles"],"_links":{"self":[{"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/posts\/14878","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/users\/11"}],"replies":[{"embeddable":true,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/comments?post=14878"}],"version-history":[{"count":1,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/posts\/14878\/revisions"}],"predecessor-version":[{"id":14881,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/posts\/14878\/revisions\/14881"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/media\/14879"}],"wp:attachment":[{"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/media?parent=14878"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/categories?post=14878"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/dtasiagroup.com\/vi\/wp-json\/wp\/v2\/tags?post=14878"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}