Speaker

Andrew Ray, Principal Data Engineer at Silicon Valley Data Science

Andrew Ray

Principal Data Engineer, Silicon Valley Data Science

Dr. Andrew Ray is a Principal Data Engineer at Silicon Valley Data Science. He enjoys working at the intersection of engineering and data science. Andrew is an active contributor to the Apache Spark project. In his past life Andrew was a Data Scientist at Walmart, where he built an analytics platform on Hadoop that integrated data from multiple retail channels using fuzzy matching and graph algorithms. Andrew also led the adoption of Spark at Walmart from proof-of-concept to production. Andrew earned his Ph.D. in Mathematics from the University of Nebraska, where he worked on extremal graph theory.

Sessions

Data Wrangling with PySpark for Data Scientists Who Know Pandas

Data scientists spend more time wrangling data than making models. Traditional tools like Pandas provide a very powerful data manipulation toolset. Transitioning to big data tools like PySpark allows one to work with much larger… Read more

Write Graph Algorithms Like a Boss

Graph-parallel algorithms such as PageRank operate on an entire graph at once. Efficient distributed implementations of these algorithms are important at scale. This session will introduce the two main abstractions for these types of algorithms:… Read more