DF
Systems, Platform and Reliablity Engineering

Dave Finster

Building, operating, breaking + fixing platforms that enable engineers, businesses and services to thrive.

Experience
2018 — Now
Staff Systems Engineer, Site Reliability Engineering · Google

Technical lead responsible for the Fabric Networking SRE product group that spans 60 SREs in multiple sites. Our purview covers all on-machine userspace software and equipment within datacenters necessary to provide high performance, reliable and software defined networking capabilities to all Google products, services and Cloud customers.

Technical Lead, Host Networking SRE

Responsible for the technical capability, reliablity and continued rapid deployment of Snap across the global fleet managed by a team of 20 SREs spread across multiple sites. Supporting weekly rollouts that touch the critical networking path for every product and customer has required robust pre-deployment test coverage coupled with realtime health monitoring, continuous canary analysis and automated degradation response. This stance has consistently removed the need for urgent human intervention when problems occur and contributes to a feedback loop that continuously improves testing coverage.

Snap + Pony Express

Snap serves as the vehicle for delivering and regularly updating complex networking capabilities without the need for disruptive kernel updates and with negligable user disruption. Pony Express provides a comparatively more efficient and reliable transport alternative to TCP/UDP tuned for Google datacenters.

Cloud Networking

Snap implements the Andromeda data plane running on both compute hosts and Titanium IPUs. Host Networking SRE provides incident response and complex debugging support to Cloud customers experiencing issues or attempting to achieve the highest possible performance from their VPC. I personally work with many of the largest GCP customers in providing both reactive and proactive support, advice and guidance on how best to run their workloads.

AI Hypercomputer - TPU and GPU

Networking and fleet management expertise for operating, maintaining, debugging and handling incidents related to the dense supercomputer-like deployments of TPU and GPU clusters for the largest GCP customers. Architected and implemented the infrastructure for deploying impact-free software updates across all platforms that enable the delivery of bug fixes, security patches and performance/feature improvements without disrupting customer workloads.

a3-mega

Delivered a performant GPU training platform to market using a mixture of Google technologies and hardware that had never before been combined to offer a3-mega in 4 months from concept to GA. This machine design consisted of 8 x H100 GPUs combined with 9 x Google Titanium IPUs and PCI-E switches to enable GPUDirect-TCPXO via a NIC-per-GPU model.

Incident Response

Primary incident responder for home teams of Host Networking and Fabric Networking SRE. I am also part of rotations that operate exclusively for handling large scale single or multiple all-of-region incidents, serving as the lead for operational response, customer/stakeholder communication or as incident commander.

2015 — 2018
Chief Technology Officer · Docketbook

Bootstrapped a startup company with the aim to fundamentally change how data is exchanged within the civil construction and associated industries. We built solutions to replace the ubiquitous, voluminous and manual proof-of-work processes with a universal format that serves as a single source of truth. These offerings eliminate data (re-)entry of the same proof across multiple companies, dramatically reduced the risk of loss/disagreement and rendered entire classes of fraud impossible. Started as a company of 4 people and continues to operate with a staff of 10.

Idea to MVP

Despite having a good understanding of the problem space, we worked and iterated with an initial set of interested customers to ensure we were building a product that served them successfully and would be accepted by their workforce in the field.

Customer Partnership and Integration

Worked directly with customers both in the field and corporate side to ensure we understood their problems, workflows and opportunities to improve. This effort led to shaping our API offerings, exposing data where analysts and engineers already work (typically Excel) and identifying automation points that allowed folks to just do their job and have the paperwork be handled for them.

Full Stack Development

As a small team, there were no barriers to which part of the stack we contributed. I designed and delivered databases stored in RethinkDB and CockroachDB accessed with backends written in Node.js or Go powering frontends written in Angular, React and React Native.

Containerized Infrastructure

Adopted containers early to streamline not only the developer experience but to provide consistency in how we deployed in our staging and production environments. Any engineer could bring up a full fidelity environment on their laptop in minutes and iterate quickly. By mimimizing divergences between local development and production-like environments, we enjoyed a high level of test reliablity and deployment success which in turn supported consistently low developer friction and high velocity.

Private and Public Cloud

Designed, built and operated both an "on-premise" private cloud deployment where we sourced our own colocation space, equipment and diverse BGP-speaking IP transit/peering feeds. As we grew and as Google Cloud entered Australia, we expanded and eventually relocated the staging and production environments into Google Kubernetes Engine.

2013 — 2015
Manager, IT and Software Engineering · Seymour Whyte Group

Responsible for a team of 5 that handled the design, operation and maintenance of all corporate IT infrastructure across a national civil constructor. We supported and enhanced the core business by delivering bespoke technology solutions backed by reliable internet connectivity in extremely remote locations cost effectively and rapidly.

IT Infrastructure

Replaced an aged infrastructure with a modern stack consisting of self-designed Supermicro servers running Joyent SmartOS connected via a newly deployed and self managed MPLS network consuming L2 wholesale links. Converged critical line-of-business and general applications into newly leased colocation space and consistently delivered application availability/performance beyond SLA. Handled all security, compliance, supplier and capability aspects of both the backend and end-user fleets.

Office in a Box

Portable and rapidly provisioned solution for standing up connectivity in arbitrary locations. By leveraging the ACMA radio license database along with data from surveyors, we were able to model the propagation of cellular coverage to identify optimal antenna selection and placement. This cellular link was coupled with a custom layer-2 networking solution to establish connectivity back to our national MPLS network and leveraged packet-level deduplicating WAN accelerators plus transparent local smart caches to provide the best experience possible. This capability reduced site mobilization timelines by months, allowing more staff to transfer sooner instead of waiting for higher bandwidth fixed-line fiber to be built.

Productivity Enhancing Bespoke Software

Developed several in-house applications to improve in-field data collection correctness, tighten feedback loops between sites and corporate HQ and reduce manual data (re-)entry. I developed and operated the iOS native apps and Node.js backends whilst other team members developed C# .NET applications.

Aquisition Infrastructure Merge

Absorbed the legacy infrastructure of an aquired company without any impact to end-user operations. This allowed us to not only improve the standard of service at the aquired company but our existing infrastructure easily accomodated their needs.

2010 — 2013
Lead Software Engineer and Architect · Construcsys

Developed mobile software for the civil construction industry whilst also providing IT consulting support for modernizing infrastructure at various companies.

Site Diary iOS

Improved the quality of data collected in the field and reducing reporting delays to project managers. Our software was written as native iOS applications for iPads with a large scale cloud component being created in Ruby running on top of Amazon Web Services. Responsible for technical requirements elicitation (extracted from functional requirements), software development and testing, database design and administration as well as infrastructure maintenance and performance tuning.

Consulting

Performed several on-request consultancies that resulted in overhauls of existing IT systems and custom ERP software being developed/deployed with data carried across from multiple sources (a particularly difficult one being Filemaker)...

2009 — 2013
iPhone and Web Developer · Internetics

Responsible for designing and implementing iOS applications, occasionally with back-end cloud/web services, from scratch for various clients. Also provided technical support where required in the form of server administration and software process development. All development was done using Objective-C on the iPhone and iPad with Ruby and PHP used elsewhere.

2010 — 2011
Software Engineer · Raytheon

Provided consulting services for the development of intelligent, self-managing IT infrastructure and trainable machines.

2008 — 2009
Systems Administrator & Software Developer · Veitch Lister Consulting

General IT administration and assisted with the roll-out of new technologies such as email services, advanced development servers, VoIP systems and multi-office networks. Created several in-house software tools for use with the newly deployed VoIP network as well as other utilities. Worked on the development of traffic simulations for the purposes of identifying road network bottlenecks and the demographics of road users passing a particular location for outdoor advertising applications

2006 — 2008
Junior Software Developer · Technology One

Developer in the Research and Development division. Worked on the frontend of Works and Assets as well as Web Services. Also assisted with writing new unit tests and converting old ones to suit new testing frameworks and changes in the software logic.

Projects
All projects
active 2026
Plex Edge Cache, Prober and Performance Monitoring Pushing the limits of how performant a homelab Plex server can be.
active 2025
TSInject Everything should have it's own Tailscale
active 2025
ConfigFS How easily can we spread out small config files to arbitrary nodes.
active 2025
Flyscale While mesh networking is nice it is easier to reason about the security stance of proxies
active 2025
Local Dev Environment While mesh networking is nice it is easier to reason about the security stance of proxies
active 2023
Homelab Because who doesn't like to bring all the problems at work home with them?
Writing
All posts
Sep 2026
One Broken Bolt