Chemical Data Scientist
WFA Digital Insight
Valdera’s Chemical Data Scientist sits at the intersection of chemistry and data engineering, a niche that few remote teams tackle at scale. The role isn’t just about writing code; it demands a deep familiarity with chemical identifiers—CAS numbers, SDS documents, and NAICS classifications—so the data you clean actually translates into actionable supplier insights. You’ll own the end‑to‑end pipeline, from scraping public catalogs to normalizing messy supplier feeds, and then feed that into matching models that power Valdera’s procurement platform. Collaboration is built‑in, with the Supplier Management and Engineering squads shaping data quality standards together. Candidates who can prove both chemical domain fluency and solid Python‑driven automation will feel right at home here.
Job Description
About Valdera: At Valdera, we empower innovators to turn ideas into reality by transforming how manufacturers source materials. We make it effortless for companies to find the best materials and suppliers for their needs, enabling them to build high-quality products at scale and deliver them to millions of consumers worldwide. We are a team of ambitious, results-driven individuals with a proven track record of working with Fortune 500 industrial manufacturers, beauty brands, and chemical companies. We are a fast-growing company that hires talented, hardworking people who excel in high-performance environments and want to grow their careers quickly. Our culture is built for exceptional individuals to take on meaningful challenges, collaborate with the top minds in our industry, and see the direct impact of their work. If you’re looking for a fast-paced environment where your ideas will drive real change, Valdera is the place for you. Join us, and let’s shape the future of manufacturing together. Role Description: We are hiring a Chemical Data Scientist to build and maintain the pipelines that keep Valdera's supplier and chemical product data accurate, current, and structured — the foundation every buyer and supplier relies on across Valdera's procurement platform. Data quality plays a critical role at Valdera. When a buyer launches a request, they expect to be matched with the right suppliers and accurate specs on the first try. Delivering that depends on clean, current data — CAS numbers, specifications, certifications, and regulatory documents pulled from thousands of inconsistent, often messy sources. This requires strong data engineering fundamentals and a working knowledge of chemical industry data. For example, you might take dozens of differently structured chemical supplier catalogs and turn them into one clean, standardized product database. You will take ownership of the full data pipeline — from scrapers and ETL workflows to data cleaning, matching, and classification models that connect suppliers to buyer requirements. You're energized by messy, real-world data and confident partnering with Supplier Management and Engineering to close coverage gaps. As a data-obsessed professional, you're dedicated to the accuracy our buyers and suppliers depend on. Role
Responsibilities
: Design and build pipelines to collect supplier data and chemical product information (specifications, CAS numbers, certifications, SDS/regulatory documents, NAICS classification of manufacturing plants) from supplier sites, distributor catalogs, trade databases, and other public and semi-structured sources Develop and maintain web scrapers and automated ETL workflows to keep supplier and product data current at scale Clean, normalize, and reconcile inconsistent supplier data into structured, standardized formats suitable for internal tools and analytics Apply chemical domain knowledge to validate and enrich data — resolving product names, CAS numbers, synonyms, and specifications across suppliers Evaluate and improve matching and classification models to map suppliers and products to buyer requirements, and to identify overlapping or equivalent chemical offerings Partner with Supplier Management and Engineering to define data quality standards, identify gaps in supplier coverage, and prioritize new data sources. Own pipeline health and data quality, and drive the KPIs that measure overall data coverage Experience &Qualifications
: 5+ years of experience in a data science, data engineering, or applied data role, ideally with exposure to messy, real-world or industrial datasets. Working knowledge of chemistry or chemical industry data — comfort with CAS numbers, chemical properties, SDS documents, NAICS classification, and supplier certifications Strong Python skills, with experience building web scrapers and data pipelines Experience with data cleaning and normalization at scale, and a good eye for spotting inconsistencies in unstructured data Familiarity with building or applying matching, deduplication, or classification models (traditional ML or LLM-based approaches) Hands-on experience using AI tools and LLMs to accelerate data extraction, enrichment, or engineering workflows Startup mindset with a strong sense of ownership — comfortable working independently in a fast-moving, remote environment with ambiguous, evolving priorities Salary Range: Salary ranges are determined by multiple factors, including the labor market, market compensation bands, internal parity, and budget considerations. The final offer will be based on the candidate’s individual skills, qualifications, location, and experience relative to the requirements of the role.Benefits
Valdera offers generous benefits to employees. You will be provided a more detailed breakdown of your options prior to joining Valdera. Equal Opportunity Employer Statement: Valdera is an equal-opportunity employer committed to building a diverse and inclusive team. We welcome applicants of all backgrounds and celebrate a culture that values varied perspectives, skills, and experiences. We are dedicated to maintaining a workplace free from discrimination, where everyone feels valued, respected, and empowered to contribute.How to Stand Out
- Showcase a portfolio of Python scrapers or ETL pipelines that handle unstructured chemical data; include code snippets or GitHub links.
- Highlight any experience you have with CAS numbers, SDS documents, or NAICS classifications in your résumé and cover letter.
- Prepare to discuss specific data‑quality challenges you’ve solved and the metrics you used to measure success.
- Demonstrate familiarity with AI‑assisted data extraction by sharing examples of prompts or LLM workflows you’ve built.
- During interviews, ask about Valdera’s data‑quality KPIs and how the team prioritizes new supplier sources.
- If offered a salary range, research typical compensation for senior data engineers in the chemical domain to negotiate confidently.
- Watch for signs that the role expects overly broad responsibilities without clear ownership; ask clarifying questions about reporting lines and success criteria.
This is a remote position listed on WFA Digital, the platform for professionals who work from anywhere. Browse more remote jobs across all categories.