About Me

I have been a data scientist since the early 2000s, before the field acquired its name. After obtaining an undergraduate degree in geology at Cambridge University in England (2000), I completed Masters and PhD degrees in Astrophysics (2001, 2004) at the University of Sussex. I then moved to North America, completing postdoctoral positions in the subject at the University of Illinois at Urbana-Champaign (2004−9, joint with the National Center for Supercomputing Applications), and the Herzberg Institute of Astrophysics in Victoria, BC, Canada (2009−2013).

Moving to the San Francisco Bay Area, I joined the startup company Skytree “The Machine Learning Company” in 2013, that was based upon the world’s fastest machine learning algorithms of the FASTLab at the Georgia Institute of Techology. In 2017 the Skytree technology and team were acquired by Infosys. A short stint at Oracle was followed by the startup Dotscience, then Paperspace in 2020. Acquired by DigitalOcean in 2023, Paperspace’s expertise in GPUs and AI formed the foundation of DigitalOcean’s successful move into the AI space from 2024 onwards.

In 2025 I left DigitalOcean to pursue geospatial data science, following a longstanding interest. I am currently working freelance in the area. Data science, GIS, and GeoAI projects are in scope, using my experience in data science and AI, and acquired skills in GIS.

Machine learning has been part of my work since 2000, first applying it to large astronomical datasets, followed by wide ranges of application at the companies mentioned above.

My approach to data science is that the most important part of the process is always the business (or other) problem, clear agreement and understanding by all interested parties of what value will be created by its solution, and what end-to-end analysis needs to be done to deliver this. Rigorous data understanding and preparation is often necessary, and high quality data should be used or created where possible. This in turn enables the power of analytics, AI and machine learning as tools to get the best result. Production-readiness, if required, is implicit in this end-to-end approach via the understanding of what is to be delivered. Usually full production requires the collaboration of engineering.

This end-to-end philosophy is independent of the AI/ML or other software tools used to instantiate it, and so remains relevant in the age of deep learning, LLMs, reasoning models, agents, and whatever comes next. It is appropriate in geospatial in a similar manner to other areas of data science.

Some highlights from my work include the following (see here for more details):

  • Data Mining and Machine Learning in Astronomy: 61 page review published in the International Journal of Modern Physics D in 2010 was arguably the first such extensive review of this now significant field.
  • Skytree PoCs: From 2013-17 Skytree never lost a proof-of-concept (PoC) customer engagement to a rival company based on performance. I led or participated in PoCs with over 60 companies, including deal sizes of $1m+, international and onsite work.
  • CANFAR+Skytree (2012): We combined the compute power of 500 nodes in the Canadian Astronomy Data Center with the algorithmic speed of Skytree to enable large-scale machine learning in astronomy, one of the first projects in the world to do so.
  • Principal Investigator on Robust Object Classification in a GALEX-SDSS Federated Dataset, a $45k funded NASA grant (2006-8); co-I on NASA grants totaling $970k (2004-9) as part of the Laboratory for Cosmological Data Mining at Illinois/NCSA.
  • Skytree benchmarking: Leveraging its underlying C++ codebase and FASTLab algorithms, I led the benchmarking project that showed Skytree to be over 100x faster than R and Scikit-Learn at the time (2015) for common ML algorithms.
  • Morphological Classification of Galaxies in the Sloan Digital Sky Survey using Artificial Neural Networks (2004): Classified at a rate of 300,000 galaxies per minute, cf. the well-known Galaxy Zoo (now Zooniverse) project that did this via citizen science over several years.
  • Trillion element training set: we constructed and ran a training dataset of this size using Skytree’s gradient-boosted decision trees on Hadoop distributed computing, a decade before modern LLMs.
  • Chief data scientist (CDS) at Dotscience: Titled Principal, this was de facto CDS, and I owned and drove the product roadmap, defining and prioritizing the most critical revenue−driving features, such as RBAC, A/B testing and enhanced statistical model monitoring capabilities.
  • Wide-ranging projects at Paperspace and Digital Ocean using GPUs and LLMs at scale.

Outside of work, I play tournament Scrabble, and enjoy reading, walking, and road trips. I am married with 2 sons, born in 2018 and 2024.

For more, see the other sections linked in the navigation bar.