This workshop represents the next step in establishing the Zephyr Foundation, a rigorous, community-wide framework for improving travel analysis methods. Attendees will work in teams to develop scopes for specific projects that improve travel analysis in order to support better decision-making and the public good.
The following is a list of the projects that were discussed at the workshop. These are a selection and combination of the many submittals that we received and only represent possible projects. The Zephyr Board will be in charge of selecting projects that the organization will pursue. Expand a project below to read the full workshop notes.
As our data sets become larger, and our models more complex, the limits of the scientific paper as the basic unit of trade are increasingly clear. What can Zephyr do to help incentivize practical reproducible research?
Scientific journals have long been central to ensuring the validity and credibility of science through data recording and sharing, reproducibility and external review. However, as our data sets become larger, and our models more complex, the limits of the scientific paper as the basic unit of trade are increasingly clear. Put simply, a complex model is difficult to describe in 10,000 words. Even in cases where a model is well-formulated and clearly expressed mathematically in an available peer-reviewed paper, it can still be a major challenge to reproduce and build upon existing research. It may be that the software or data are sitting on a lone researcher’s computer, or the code may be impenetrable to anyone other than the original author. Or it may be that the model is carefully crafted to work in one specific instance, so applying the model to a different city is challenging. What can Zephyr do to help incentivize reproducible research?
Important question: is the journal model getting obsolete? Maybe we should be creating a new venue for sharing.
Distinction to be made: Behavioral research vs. implementation (the way you explore the behavioral thing is by testing the model - a lot of behavioral assumptions tend to be built into models).
Enables hypothesis to be tested, which facilitates verifiable progress. This is hard to do strictly from the words of a scientific paper.
Once you publish something as an innovation, we should have an expectation in the field that it should be tested. Lack of reproducibility allows singular results, which could be clouded by enthusiasm bias/commercial interests, data anomalies, and errors to move forward as “truths” as opposed to a meta analysis of multiple studies. In other scientific fields, another researcher often must replicate results.
If research is made more reproducible, the hope is that it will generate more useful results to be used in practice.
NDAs and human subject requirements — Suggestions: “scrubbed” datasets, remote computing, easy to reference NDA usage templates. Learning from other fields: medical records.
Complex systems don’t easily lend themselves to being described in paragraph form — Large model systems are hard to describe in a singular research paper where space needs to be reserved for “the innovation,” and impossible to build replications of based on a word-for-word description. Suggestion: well-documented code — if there was code, you could dig in; not a small task but possible; code needs to be documented and have metadata.
Innovators aren’t necessarily the best communicators…or coders — Suggestion: need to value the communication as much as the innovation.
Need economic incentives, but proprietary methods impossible to validate — Suggestion: standardized validation datasets for benchmarking.
Before and after studies are useful to analysis and model developers as well as planners and policy-makers to learn from mistakes, calibrate tools, and bookend probable outcomes. What could Zephyr do to support systematic before/after studies in our industry?
Analysts and model developers can use these studies to learn from mistakes and provide data to calibrate their tools. Planners and policy-makers can use them to bookend probable outcomes, learn from mistakes, and point to other successes as examples. However, most existing efforts to systematically document before and after states have focused on a specific type of forecast (major transit projects or toll roads) or a specific agency, which makes it difficult to draw comparisons between types of models and methods. Zephyr could support before-and-after data collection efforts and analysis that would best occur before construction/implementation of a significant transportation policy and project, and then again after regular operations began (similar to how FTA’s New Starts program works). Zephyr could also support, maintain, and manage the data storage perpetually rather than on a project-by-project basis.
Forecasters: improved processes; better reputation; more support from data.
Funding that could take away from project budget; self-interest (don’t want to look bad); once a project is “done,” there is little interest in re-opening it.
Previous work on before/after studies conducted by FTA showed that almost all models have flaws, are inaccurate, and unreliable. However, this finding must be weighed against the lack of data for validation.
Objectives: improve forecast accuracy by advancing the state of the art through understanding what is happening today; provide accountability; estimation of project benefit.
Develop “best practices” for before/after studies and the ongoing programs that support them (there are likely already some good examples, e.g. the World Bank). Types of practices could include: data collection (ongoing, and at start/end of project), data standards, sensitivity tests for inputs, data and analysis tool storage, performance metrics, appropriate timeframe for evaluation and period of performance of facility, features of the evaluator, and costs and funding strategies. Best practices will likely differ based on the type of study or needs placed on specific forecasts. Consideration should also be given to the variety of forecasts often produced behind the scenes to develop the project — these often involve multiple sensitivity tests of potential scenarios and are not always published, since most required documentation requests a single point forecast or a specified set of inputs that may not be reasonable.
Types of performance measures: the original scope; predicted and eventual capital cost; transit service level; project ridership. Conflicts of interest should be appropriately considered — for this reason, academics or other third parties may be the best “evaluators.”
Execute specific before/after studies — could be used to test the best practices. Evaluate a specific “shovel-ready project” like a managed facility.
There are probably a lot of potential funders and interested parties that could be willing to pay for such an effort, which should push this up the priority list. However, these studies are not inexpensive, and it could be difficult to motivate an agency that has already “closed its books” on a specific project to pay for something additional.
Well-managed and designed community-owned software projects could be broadly applicable to a number of public needs, obviating unnecessary duplication. Unfortunately many open-source projects lack a permanent and actively managed home. Would Zephyr be a good home for such projects?
The travel analysis community has developed and contributed to a multitude of open-source software projects, motivated for a variety of reasons and often either complementary to commercial analysis products or fulfilling a need that is not yet commercially viable. A sample of the types of open-source travel analysis software under active development at the time: travel modeling frameworks in use by agencies (ActivitySim, DaySim, CT-RAMP, MATSIM, Tranus, etc.), dynamic network assignment (Dynus-T, DTALite, Fast-Trips), sketch planning packages (the VisionEval family), transit assignment (Fast-Trips), land use models (UrbanSim), and supporting packages and tools (benefit-cost tools, ORCA, Pandanas, Wrangler, CountDracula, dtaAnyway, helper scripts, population synthesizers). These projects are wide-ranging in scope, management, standards, and ownership. A common theme: if well managed and designed, they could be broadly applicable to a number of public needs, obviating unnecessary duplication. Unfortunately many of these projects lack a permanent and actively managed home due to projects and funding ending, or public agencies being necessarily focused on their specific needs rather than designing and maintaining for a broader audience. Would Zephyr be a good home for projects such as these, and provide (with assistance from others) the necessary administrative, financial, and technical resources? Should Zephyr also create a directory of open-source projects and code so that one could find what already exists?
There are a variety of software products that accomplish similar things around the world, and a lack of coordination. Code is often written for the specific needs of a single agency rather than being built for transferability. Agency-owned code is often not well-maintained — when it is, different agencies pay for consultants to fix their specific code, when it could be easily consolidated. Why not have a piece of software that we are all constantly improving, instead of fixing when hired for a specific project?
Value propositions: exposure, higher software quality, efficiency of money and time.
Endorse a “toolbelt” so that as much as possible we can focus our efforts on knowing the same tools: preferred languages, preferred data standards, preferred support packages (e.g. pandas), preferred testing/integration suites. Considerations: choices should be accessible for an average student; what can more people contribute to; standards are about 25 people agreeing with something, not someone saying they’re going to write a standard; the “best” is the one that is mostly used — usage will say what’s more relevant.
Develop a funding mechanism — a Software “Membership Fee” (each agency puts a fixed amount into the foundation, and a % goes to improvement and maintenance of a suite of “Zephyr Projects,” including adding/improving documentation and simplifying/improving code), or a Project “Membership Fee” (each agency gives a fixed amount to a specific project built with the pooled money). Consideration: a chicken-and-egg situation — if we had a good project already established we could use it to attract funding, but if we have funding we can do the cool projects. How to start them?
Write rules — have a clear contract to both contributors and clients.
Ensure quality — conduct testing of the software to guarantee it works in different situations, with a quality control group. Considerations: after 10 years it will need to be re-written and updated — who will do that? Extensibility and use of common data standards and APIs should be encouraged.
Manage development and maintenance of specific products — Zephyr could issue contracts; agencies would give money to Zephyr to maintain the different/common software, and Zephyr would contract with consultants. Zephyr should own the repository, organize and maintain it, and add contributions. How transferable are models and who owns the model? If a specific entity owns it, efforts can be optimized.
Structure project management — develop a clear management structure to keep, maintain, and update Zephyr software projects, and define the required resources (monetary and human). Considerations: for the project to survive you need the owner committed to keep driving it; an incubated project may be different than an established one.
Zephyr should keep software in different stages of development, having both an incubator and a platform for more developed software. Considerations: what would be easiest to succeed — the smallest or the biggest? Should have a history of being open source, robust code, easily extensible (no single giant class), and enough users. We cannot put all efforts into a single model. ActivitySim could be a good starting point, as well as OMX APIs.
How do we consider impending transformational changes within the span of our forecasts that are clearly outside the bounds of elasticities we can currently observe? This project considers “uncertain topics” such as the introduction of automated and connected vehicles and discusses how Zephyr could help.
Current travel demand models are observed-conditions models. For the first time in decades, the industry is faced with transformational changes within the span of our forecasts that render those conditions obsolete. This proposal is to create a parallel version of a model which considers “uncertain topics” such as the introduction of automated and connected vehicles. The model would exist as a set of alternate procedures that could be toggled on and off, and would inform certain changes to the structure of the observed-conditions model necessary to invoke these changes. What would such a model need to consider? Changes to highway operations such as capacities, speeds, volume/delay functions, signal operations; changes to how travel costs are experienced or perceived; changes in how passenger vehicles are allocated and used; and more. How could Zephyr help spearhead this work?
Decision-makers need both answers they can use and a full understanding of the spectrum of likelihoods. How can policymakers organize themselves to maximize their impact in affecting technology development that is favorable for society? In many cases policymakers are already skeptical of the results of these models, so we should seek to minimize their uncertainty.
Most forecasters make educated guesses about uncertain futures and present ranges such as conservative, moderate, and liberal. But spectrums are hard to understand, hard to create, and in many cases not allowed by existing political and legal processes.
Examine different ways of using existing models — document a different way of using existing models to forecast uncertain topics, emphasizing scenario planning while retaining the trust of decision-makers. Present model results to decision-makers in a way that lets them test different scenarios and alternative futures. Make degrees of freedom easier to tweak in order to more easily do scenario analysis / reduce embedded assumptions.
Synthesize, critique, and disseminate existing research and practices — a considerable amount of people are attempting to model CAV behavior, with lots of research being conducted as well. A synthesis and critique of research and forecasting approaches would be helpful.
Provide recommendations for updating existing models — identify parts of existing models (demand and network) that are likely to require sensitivities, and recommend elasticities to test. Value of time and auto availability models are obvious examples.
Develop a new forecasting paradigm — engage decision-makers and modelers in developing a new forecasting paradigm involving “futures” rather than “forecasts.” Zephyr could bridge the discussion so the two groups can identify weaknesses in models and brainstorm ways to make the models more extensive. Zephyr should engage in advocacy for allowing models to appropriately represent uncertainty.
Explore new model types — models are traditionally based on observed behavior. Do we need a change in the entire paradigm of modeling to forecast these uncertain topics? Maybe Bayesian techniques?
This team explored what it would take to rigorously assess model performance in a testbed environment and determine what resources would be needed to make it successful.
Zephyr could consider providing research grants to academics and other researchers to calibrate their model systems for a single metropolitan area, then assess the models’ performance across time — ideally a period of ten years — to see how well they perform. This project would explore what it would take to do this in a rigorous and useful way; how to select the test bed and identify the set of stakeholders; and determine what resources would be needed to make it successful.
Help regions identify the models and tools that best serve their needs, and save regions the trouble of producing the same thing over and over (particularly if it isn’t useful). There is likely value in evaluating sketch planning tools vs. complex models; mixed thoughts on whether to include land use models.
Identify test beds — should there be multiple test beds with different features/sizes? It’s a real challenge to find a singular “representative” location. Suggested comparisons of 2010 and 2020. Potential criteria for determining where a test bed should be: what “evaluatable” changes happened over the 10-year period vs. what sensitivities do we want to test; willingness and motivation to participate; political resilience.
Define performance evaluation process — requirements for doing a comparison include being calibrated to the same data and achieving similar validation. Definition of a “better” model system depends on the questions being asked, and will likely be very different for small cities compared to complex regions. Metric types include screenlines, VMT, other regional performance measures, and project-level metrics.
There were a diversity of opinions, but in general the group thought this was a mid-level priority. There are other topics that would need to feed into this one.
Travel analysis will increasingly rely on massive amounts of data collected using imperfect and inconsistent methods. Yet data wrangling can easily eat up a lot of bandwidth or budget. This team explored whether it would be helpful for Zephyr to incubate an open-source data wrangling interface.
Agencies and consultants that have the bandwidth for utilizing these data sources undertake a significant amount of start-up time wrangling the data, while others rely on either data providers (who typically do not divulge their methods) or consultants (who often are doing the same thing over and over again, reducing efficiency). Some agencies like SFCTA have developed nascent open-source tools to store and wrangle data and interface with travel models, such as “CountDracula.” However, public agencies are not good owners of open-source products. This project would (A) create a standardized data schema for observed travel data; (B) build upon the CountDracula data management tool; (C) extend CountDracula to add a “Count Dracula’s Lab” which would add data fusion and cleaning features; and (D) extend CountDracula’s visualization features to be more public-facing.
Variety of data types: origin-destination (person flows); point (flows, crash); network (flows, projects); area (event, weather, land use). Data schema should have the capability to tag.
Survey agencies to identify state of the practice — what they collect, current process, storage, how they use it, what could be done better, freshness policies, concerns of privacy, and (IT-specific) privacy, data storage, and which databases/software are used.
Identify and organize interested agencies and implementation partners — big data providers, smartphone apps, enthusiastic agencies, academics, and concrete working groups with academics/industry/agencies. Agencies could contribute working hours/staff instead of funding; data providers could contribute data instead of funding; could brainstorm and produce proof of concepts for various features using a hackathon.
This team explored what it would take for Zephyr to support a standardized network format.
Much like Google standardized the data format for current (and archived) transit networks, the travel modeling community needs a standardized way of coding, using, and sharing travel network scenarios — past, current, and future. Generalized network specifications would reduce staff training time, scenario development time, and consultant costs. This group roadmapped what a Zephyr-supported data standard development process would look like, using a general travel network specification as a specific example.
The inordinate amount of time required to develop the network by MPOs needs to be resolved. Currently there are too many different solutions — inelegant, non-transferable solutions to the same problem.
Convene stakeholders to identify requirements — a data model with the capability of storing current, past, and future conditions (including variants on the future), and good capabilities for switching between scenarios. Should have versioning abilities at very least, which can evolve to scenarios. Must be sharing-oriented, which could greatly augment the ability to visualize results. What resolution will it support — micro vs. macro? Dynamic? What schema — node-link? intersection geometry? signalization?
Identify existing solutions — OpenStreetMap; the Jellyfish project (Queensland University attempted a master network for Australia). Examine where the time-consuming steps are in current workflows.
Standardize network vocabulary — decide on standard naming conventions; focus on static assignment as a first pass; leverage existing popular standards where possible (e.g. GTFS, OSM). Need some sort of “wow” factor for phase I, which this currently lacks.
Standardize network schema — decide on directed or reversible links, a minimum set of schema fields, and a set of 3-5 data formats that Zephyr can support APIs for and create data validators to check network topology. Reference research on existing formats and solutions.
Develop a network management strategy — capabilities for diff-able networks, scenario development, version control/flow standards, and projects/master networks.
Involvement should include vendors, consultants, and owners — make sure there is a variety of players to ensure different ways of thinking.
We believe this could be a quick win.
Effective travel analysis, planning, and policy-making requires credible and useful data that can be efficiently exchanged, analyzed, and validated. This group explored why and how Zephyr and professional organizations can develop data standards and guidelines.
The status quo does not support standardized measures of data validity, data management and maintenance, data schema, or dissemination. The inconsistency of data across geographies, jurisdictions, and companies creates significant costs to researchers, analysts, and data consumers that could be remedied with market-supported data standards and guidelines.
Examples of recent headaches with data management: using observed traffic signal data to predict what an operator would do, with no consistently formatted data available; merging data from across states; confidentiality when collecting data (hard to guarantee confidentiality with the Feds reading emails); getting a data server set up is challenging; questionable quality of household travel survey data for a research project; anonymizing schemes for fare card data met lots of legal challenges; connecting call record data with other data sources raises privacy/anonymization challenges; survey response rates.
Fellowship program — addresses workforce issues; a Data Officer could go into a public agency, similar to how Code for America works.
Standardize vocabulary and data dictionary — a lot of time is wasted dealing with different variable names and terminology; Zephyr could set standard variable names and units (e.g. commuter_rail vs. CommuterRail).
Standardize schema — there should be a periodic update of standards so they don’t become stale or forked. The Zephyr board could decide each year whether to revisit each standard; a technical standing body would monitor data standards to inform the board.
Standardize metadata and data projects — a standard protocol for managing data projects or developing data resources, including who is responsible, a data dictionary (specific field names & types) for various types of data that grows and evolves over time, metadata standards, and updating/sunsetting.
Develop example contract language — Zephyr can help come up with example contract and licensing language to facilitate data sharing with researchers and other stakeholders while protecting privacy.
Create a how-to manual for collecting travel analysis data, similar to a survey methods manual but for other data sources.
Data validation and grading system — how to grade a dataset: completion, sample size, representation (a dataset grade would depend on the specific application). Potential actions: develop scripts to analyze suitability and assign a grade to datasets; award datasets badges by use type; let the private sector grade its own data using the scripts, or let Zephyr act as a “ratings agency” with access to evaluate; organize dataset ratings and previous uses.
For data sets unique to transportation, Zephyr can set standards but needs to talk to public agencies and groups that collect data, as well as private data providers. After that relationship is established, both groups work together to get the right data — preventing each public agency from having to independently examine each private company’s data.