Role overview
About this role
Come forge the future like an IBMer. Ready to think boldly, work with some of the world's most recognized brands, and kick-start your career? Welcome to IBM's Associate Program for university hires. From day one, you will collaborate with global clients and IBM teams on projects that help organizations solve tough challenges across digital transformation, cloud strategy, AI adoption, process redesign, analytics, modern data platforms, and agentic AI-enabled data transformation. As an Associate, you will work alongside a global cohort of diverse, ambitious peers and have access to industry-recognized certifications, digital badges, and a minimum of 40 hours of structured learning per year on IBM's AI-driven learning platform, supported by coaches, mentors, and professional communities across practices. IBM's culture of internal mobility means you can explore new technologies, industries, and career paths as your interests evolve. Bring your curiosity. Grow your skills. Build what's next - like an IBMer. This role is a strong fit for builders who like turning messy data into reliable systems and who want to help make analytics, generative AI, and AI agents useful, reliable, and safe through governed data, trustworthy context, observable pipelines, and secure interfaces to enterprise systems. To give yourself the best opportunity for success, we advise applying only to roles that align with your skills and experience, rather than applying broadly across all entry-level positions. You'll receive a status update email for each application, so be sure to check your IBM Careers account regularly — it's the best way to get a centralized view of which roles you have active applications against. As an Associate Data Engineer, you will help design, build, and improve data platforms, data products, and services that support analytics, machine learning, generative AI, and agentic AI solutions. You will work across data gathering, ingestion, transformation, storage, batch and real-time processing, semantic enrichment, retrieval, APIs, visualization, data quality, observability, and governance. Collaborating closely with diverse teams, you will play an important role in selecting suitable data management systems and identifying the critical data needed for insightful analysis. As a Data Engineer, you will help tackle challenges related to database integration, data quality, and complex structured and unstructured datasets. Key Responsibilities may include: Assisting designing and implementing scalable data architecture, data products, and management systems for modern cloud environments, analytics, and AI-enabled use cases. Work on optimizing existing data pipelines, retrieval indexes, and data services for improved performance, reliability, data quality, and freshness of AI-ready context. Collect, prepare, and analyze structured, semi-structured, and unstructured data to identify trends, providing clients with actionable insights that enhance marketing, operational, and business practices. Participate in troubleshooting data-related issues, working to solve data quality challenges, retrieval-quality gaps, processing failures, and inconsistencies affecting analytics, generative AI, or agentic workflows. Create visually compelling and user-friendly dashboards, reports, and observability views to communicate findings, pipeline health, and AI system insights to both technical and non-technical stakeholders. Ensure data integrity, accuracy, reliability, lineage, and access control through rigorous data cleaning, validation, preprocessing, cataloging, and governance practices. Work with project teams to prioritize and translate client requirements into current and future operational scenarios, processes, models, use cases, data products, APIs, tool interfaces, plans, and solutions; collaborate with clients, architects, and AI engineers. Present analytical findings, data quality insights, and recommendations clearly and concisely, demonstrating the value of data-driven and AI-enabled decision-making to clients. Work with cross-functional teams to tackle complex business problems, utilizing data expertise across cloud platforms, RAG, vector search, APIs, responsible AI, and agent orchestration patterns while staying current on modern data stack trends. Consulting And Collaboration Skills Analyze business processes and application portfolios to identify where teams need better data, context, tools, automation, or process optimization. Use Agile ways of working, planning, and project-management practices to deliver production-ready data and AI solutions. Communicate with curiosity, clarity, and empathy; ask insightful questions about data meaning, risk, governance, and business impact. High School Diploma/GED. Familiarity with one or more programming or query languages such as Python, SQL, Java, Scala, or JavaScript. Foundational understanding of data engineering concepts, including data pipelines, databases, APIs, distributed processing, ETL/ELT, data modeling, or data products. Basic understanding of cloud computing environments such as AWS, Azure, Google Cloud, IBM Cloud, or similar platforms. Ability to apply foundational statistical, machine learning, or information retrieval concepts to data preparation, analysis, search, or model-support work. Interest in AI/ML, generative AI, agentic AI, intelligent automation, RAG, LLM-powered systems, or enterprise AI platforms. Strong analytical thinking, problem solving, adaptability, teamwork, and communication skills. Willingness to travel up to 100%, based on project requirements. Bachelor’s degree in a related field such as Computer Science, Data Science, Statistics, Mathematics, MIS, Engineering, AI/ML, or another quantitative field. Coursework, projects, internship experience, or portfolio work involving data engineering, software engineering, cloud platforms, analytics, AI/ML, or modern data products. Experience with technologies such as Spark, Hadoop, Kafka, Airflow, dbt, Databricks, Snowflake, Delta Lake, Linux, or similar platforms. Familiarity with LLMs, embeddings, vector databases such as Pinecone, Weaviate, orpgvector, retrieval systems, prompt workflows, RAG, and AI agent architectures. Exposure to orchestration and agentic AI frameworks such asLangChain,LangGraph,LlamaIndex, Semantic Kernel, AutoGen, MCP-based tooling, or similar tools. Understanding of Git, containers, APIs, testing frameworks, CI/CD, Kubernetes, observability, data governance, privacy, security, or responsible AI practices. Preferred certifications, coursework, or credentials such asSnowProCore, Google Associate Cloud Engineer, Google Data Engineer coursework, or related cloud/data credentials. Familiarity with AI coding tools and GenAI concepts, including prompt engineering, RAG, fine-tuning, and model evaluation. United States Data & Analytics Entry Level San Francisco, US (0147) International Business Machines Corporation