Data Governance and Data Quality in HR.




Overwiev


Industry Category: Life Sciences & Healthcare
Technologies Used: Informatica Integration Cloud Services (IICS), Amazon Web Services (AWS), Azure DevOps, Python, Power BI.



Background


Merging two companies with different data platforms and business models presented significant challenges:


  • Data Quality Standards During Transformation: Ensuring data quality amid new employees, changing roles, and ongoing data maintenance.
  • Data Accessibility and Silos: Difficulty accessing files stored in SharePoint for analytics, leading to delays and inaccuracies.
  • Scalability Issues: Existing data processes couldn't handle increasing volumes and diversity.
  • Data Availability: Inability to perform timely analytics and reporting due to integration limitations.


Objectives


The main goal was to integrate HR files from SharePoint into Amazon Redshift using IICS, enabling analytics and reporting while ensuring data consistency and quality. Key objectives included:


  • Efficient Data Ingestion:
  • Handle various file formats and ensure efficient loading of high-volume files.
  • Implement file naming and versioning conventions for automated ingestion.
  • Enhanced Data Quality:
  • Utilize IICS's data quality tools for profiling, cleansing, and validation.
  • Apply complex data quality rules to reduce errors and improve integrity.
  • Scalability: Leverage IICS and AWS capabilities for scalable data integration.
  • Detection of Corrupt Data: Configure IICS to monitor data quality metrics and trigger alerts for failed data.


Solution


We implemented an end-to-end data pipeline integrating IICS, AWS S3, Amazon Redshift, and Power BI:


  • Infrastructure Implementation:
  • Extracted data from SharePoint, processed it through IICS, and loaded it into Amazon Redshift.
  • Created Parquet files in AWS S3 for efficient storage.
  • Built a Power BI dashboard for reporting and data stewardship.
  • Data Integration and Quality:
  • Developed mappings and transformations in IICS to cleanse and enrich data.
  • Implemented data quality rules allowing business users to manage them easily.
  • Applied validations for accuracy, completeness, consistency, and uniqueness.
  • Customization and Integration:
  • Leveraged platforms familiar to the client (IICS, AWS, Power BI).
  • Tailored transformation logic to the client's data structures.
  • Employed source control using Azure DevOps.


Challenges


  • Technical Challenges:
  • Data Profiling Issues: Profiling large data volumes initially led to long runtimes and failed jobs; we improved the success rate by optimizing infrastructure sizing, tuning IICS configurations, and running profiling incrementally rather than on full datasets.
  • Heterogeneous Source Files: HR files in SharePoint arrived in varying formats, structures, and naming patterns, often maintained manually by different teams. Establishing strict file naming and versioning conventions was essential to make automated ingestion reliable.
  • Post-Merger Data Inconsistencies: The two legacy organizations used different definitions, codes, and structures for the same HR concepts (roles, org units, employment types), requiring careful mapping and harmonization logic before data could be meaningfully combined.
  • Sensitive Data Handling: As HR data contains personal and confidential information, access controls and data handling had to be designed carefully across SharePoint, S3, Redshift, and Power BI to ensure only authorized users could view sensitive records.



  • Operational Challenges:
  • Communication and Alignment: With requirements evolving throughout the merger, we established regular touchpoints between technical teams, HR stakeholders, and data stewards to keep scope, priorities, and timelines aligned.
  • Evolving Ownership and Roles: New employees and shifting responsibilities meant data ownership was not always clear; defining stewardship roles early helped keep quality rules maintained and issues resolved quickly.
  • Change Management and Adoption: Business users needed to trust and adopt the new pipeline and dashboard over familiar manual processes, which we addressed through training sessions and by letting users manage data quality rules themselves.


Results


  • Improved Data Quality Metrics:
  • Reduced new failed records by 30%.
  • Visualized improvements through a Power BI dashboard.
  • Enhanced Reporting Capabilities:
  • Replaced the legacy reporting tool with the Global Data Governance Dashboard.
  • Reduced manual efforts in tracking and managing data quality.
  • Positive Client Feedback:
  • High satisfaction with quick access to clean data.
  • Training sessions empowered team members to use new tools effectively.


Lessons Learned


  • Effective Requirement Management: Breaking down requirements into manageable tasks with clear acceptance criteria.
  • Collaboration Across Teams: Maintaining open communication to ensure mutual understanding of scope and timelines.
  • Adaptability:
  •  Managing changes in data requirements and integration processes.


Conclusion


The project significantly enhanced the organization's data management by implementing a scalable, end-to-end data pipeline. Key achievements included:


  • Improved Data Quality: Automated data quality rules reduced errors by over 30%.
  • Cost Efficiency: Automation decreased time spent on manual data cleaning.
  • Scalability: Designed a solution adaptable to future needs using IICS and AWS.
  • Actionable Insights: Provided data quality metrics supporting data stewardship initiatives.
  • User Adoption: Empowered team members through training, leading to effective use of new tools.

Next Steps: Building on this success, we will participate in Master Data Management and Data Quality projects for other domains like Customers and Vendors, supporting both operational and development tasks.